PDF合併工具

PDF 合併 / PDF Merger

在瀏覽器內合併,不會把 PDF 上傳到伺服器。檔案很大/很多頁時會比較吃記憶體。
未選擇任何檔案
拖曳調整順序(上 → 下 = 先 → 後):

    PDF合併工具說明

    免費 PDF 合併工具(免上傳、拖曳排序、快速下載)Browser-based
    這是一個線上 PDF 合併工具,可直接在瀏覽器內將多個 PDF 合併成一個檔案:不需安裝、免註冊、免費,並支援拖曳排序與自訂輸出檔名。
    主要特色
    •免上傳、隱私友善:所有處理在你的瀏覽器本機完成,不會把 PDF 上傳到伺服器。
    •拖曳排序:可直接拖拉清單調整順序(上 → 下 = 先 → 後)。
    •自訂檔名:輸出檔名可自由命名(預設 merged.pdf)。
    •適用情境:合併報告、論文附件、合約、掃描文件、教學講義等。
    使用方式(30 秒完成)
    1
    在上方工具選取多個 PDF 檔案。
    2
    拖曳清單,調整合併順序。
    3
    (可選)輸入輸出檔名。
    4
    按「合併下載」,取得合併後 PDF。

    Free PDF Merger (No Upload, Drag to Reorder, Instant Download)
    A free online PDF merge tool that runs in your browser: no installation, no sign-up, drag-and-drop ordering, and a custom output filename.
    Key features
    •No upload: processed locally in your browser (privacy-friendly).
    •Drag to reorder: top → bottom = first → last.
    •Custom filename: default merged.pdf.

    2026年10月2日 星期五

    Duetkifu: A Kifu for Research, Recording Every Move You and Your AI Agent Make

    Working with an AI agent, you can try a dozen things a day. By the time you write the paper, only the one that worked is left. The failed attempts and the corrected numbers are scattered across chat logs.

    Duetkifu is an open-source (MIT) tool I built for exactly this. It keeps the research record you and your AI agent build together: every move, dead ends included, and where each number comes from. It works with Claude Code, Codex, Gemini CLI, Cursor, or any AI agent that can edit JSON files.

    GitHub: https://github.com/Ashur5457/duetkifu  |  Tutorial: docs/tutorial.md  |  Summary for LLMs: llms.txt

    Why record dead ends?

    A paper shows the path that worked. The rest, such as failed attempts, numbers that were later corrected and clues noticed too late, ends up in chat logs, slides and memory, and then it is gone. The faster the agent tries things, the faster this hidden part grows, and nobody can keep up by writing it down by hand.

    What Duetkifu does

    • Every attempt is a "move". Each move records which move it follows, why it was made, what came of it and, for a dead end, why it ended.
    • Dead ends are results. A failed move needs a reason and a cause: the idea was wrong, the data cannot be trusted (so the path was never really tested), the analysis method, or the cost. The next person, or the next agent, can then tell a closed path from an untested one.
    • Every number has a source. It is recomputed from raw files by recorded scripts, or taken from a named file with a SHA-256 fingerprint, or marked "no data" with the reason. Mismatches are listed, never silently rewritten.
    • The agent writes, you decide. The agent opens and closes moves and drafts their text. The page marks what the agent wrote and you have not read yet. The outcome of a move, the confirmed cause and the review marks stay yours.
    • Everything in one HTML page. The record (kifu.json) is drawn as a tree you read and edit in the browser. It also includes Duetsheet, an interactive report you review with comments and free-hand regions drawn on charts.

    About the name: a kifu is the record of a shogi or go game, move by move. Players replay it afterwards to see where the game turned and which moves were mistakes. In a duet, two play from the same sheet: a researcher and an AI agent.

    Demo GIFs

    1. Reading the research record as a tree

    Duetkifu demo GIF: clicking through the moves of a research-record tree in the browser; each move shows why it was made, its result and the data behind it

    What the GIF shows: the research record of the demo project drawn as a tree. Clicking a move opens its reason, result and data chain. The shape of a move tells how it ended (holds, dead end, correction, independent audit, not resolved, planned).

    2. Ask the AI agent while you read

    Duetkifu demo GIF: drawing a free-hand region around low scores in a chart, choosing a question tag, sending it to the AI agent, and receiving the answer

    What the GIF shows: draw around part of a figure, pick a question such as "Where does the data come from?", "How was it computed?" or "Can it be trusted?", and send it. The agent answers in a floating panel.

    3. Arrange the page

    Duetkifu demo GIF: dragging, folding, resizing and full-screening the blocks of the page

    What the GIF shows: blocks (question, tree, move, main path) are dragged, folded to one row, resized or shown full screen.

    4. Edit a move, then undo and redo

    Duetkifu demo GIF: changing the cause of a dead-end move, then Undo and Redo

    What the GIF shows: changing the cause of a dead end, then taking the change back with Undo and bringing it back with Redo. Every edit is recorded with who made it and when.

    Quick start

    You need Python 3.8 or later. It uses only the standard library, so there is nothing else to install.

    With Claude Code (terminal, VS Code or JetBrains), install once:

    /plugin marketplace add Ashur5457/duetkifu
    /plugin install duetkifu@duetkifu

    Then ask Claude to start Duetkifu on your data folder (the /duetkifu skill).

    Without a plugin, run it directly:

    python duetkifu.py "path/to/your/data/folder"

    The page opens in your browser, connected to the folder. Raw data is only read; the record and report are stored in a duetkifu/ subfolder. The interface is available in English, Traditional Chinese, Simplified Chinese, Japanese, Korean and Spanish.

    FAQ

    What is Duetkifu? An open-source research record and report reviewer for researchers working with AI agents. It keeps every move, dead ends included, with where each number comes from.

    Which AI agents does it work with? Claude Code, Codex, Gemini CLI, Cursor, or any agent that can edit JSON files (see AGENTS.md).

    Is it free? Yes, MIT license. It is a v0.8 prototype, so feedback is welcome.

    How is it different from a lab notebook or chat history? Moves are structured (parent, reason, result, cause of failure), numbers are checked against recomputation and file fingerprints, and the whole path is drawn as an editable tree.

    If you try it and something breaks or is missing, open an issue on GitHub or leave a comment here.

    Keywords: research record, research log, lab notebook for AI agents, dead ends, negative results, reproducibility, provenance, data lineage, human-AI collaboration, Claude Code plugin, Codex, Gemini CLI, kifu, Duetkifu, Duetsheet.

    2026年10月1日 星期四

    Duetkifu:記錄研究每一步的 AI 協作棋譜

    跟 AI agent 一起做研究,一天可以試十幾件事。但是到了寫論文的時候,留下來的只有「成功的那一條」,失敗的嘗試、後來更正的數字,全都散在聊天紀錄裡。

    Duetkifu(我叫它「研究棋譜」)就是為了這件事做的開源工具(MIT):保存你和 AI agent 一起做研究的每一步(move),包含走不通的死路,並且記下每個數字是從哪裡來的。支援 Claude Code、Codex、Gemini CLI、Cursor,或任何能編輯 JSON 檔的 AI agent。

    GitHub:https://github.com/Ashur5457/duetkifu  |  教學:docs/tutorial.md  |  給 LLM 的摘要:llms.txt

    為什麼要記死路?

    論文只呈現走通的那條路。其他的過程,像是失敗的嘗試、後來被更正的數字、太晚才注意到的線索,最後都留在聊天紀錄、投影片和記憶裡,然後就不見了。AI agent 試得越快,這些看不見的部分就累積得越快,靠手寫根本追不上。

    Duetkifu 做了什麼

    • 每次嘗試是一個「move」。記錄它接在哪一步之後、為什麼做、結果如何;如果是死路,還要寫為什麼走不通。
    • 死路也是研究結果。失敗的 move 一定要有理由和原因類型:想法本身錯了、資料不可信(這條路其實沒被真正檢驗過)、分析方法有問題,或是成本太高。下一個人(或下一個 agent)才分得出「真的走不通」和「根本沒測過」。
    • 每個數字都有出處。由記錄在案的腳本從原始檔重新計算,或取自有 SHA-256 指紋的指定檔案,或標示「無資料」並說明原因。對不上的地方會列出來,不會被悄悄改掉。
    • agent 負責寫,你負責決定。agent 開啟、關閉 move 並起草文字;頁面會標示「agent 寫的、你還沒看過的」內容。move 的結論、確認的原因、審閱標記永遠由你決定。
    • 全部在一個 HTML 頁面裡。研究紀錄(kifu.json)畫成可在瀏覽器中閱讀、編輯的樹狀圖;同時內含 Duetsheet,也就是可以留言、可以在圖上徒手圈選、和 agent 一起修改的互動式報告。

    名字的由來:kifu(棋譜)是將棋、圍棋一手一手的對局紀錄,棋手會在賽後復盤,看局勢是在哪裡轉折、哪一手是失誤。duet 是二重奏,研究者和 AI agent 看同一份譜。

    示範 GIF

    1. 把研究紀錄當成樹狀圖閱讀

    Duetkifu 示範 GIF:在瀏覽器中逐一點選研究紀錄樹狀圖的 move,顯示每一步的理由、結果與背後的資料

    GIF 內容:示範專案的研究紀錄畫成樹狀圖。點選一個 move 會展開它的理由、結果和資料鏈。move 的形狀表示結局(成立、死路、更正、獨立查核、未解決、規劃中)。

    2. 邊讀邊問 AI agent

    Duetkifu 示範 GIF:在圖表低分區域徒手圈選、選擇問題標籤、送給 AI agent 並收到回答

    GIF 內容:在圖上圈出一塊區域,選擇問題(例如「資料從哪裡來?」「怎麼算的?」「可信嗎?」),按一下送出,agent 的回答會出現在浮動面板。

    3. 自由排版頁面

    Duetkifu 示範 GIF:拖曳、摺疊、調整大小與全螢幕顯示頁面區塊

    GIF 內容:問題、樹狀圖、move、主線等區塊可拖曳、摺成一列、調整大小或全螢幕。

    4. 編輯 move,復原與重做

    Duetkifu 示範 GIF:修改死路 move 的原因,然後復原與重做

    GIF 內容:修改一個死路的原因,再用 Undo 取消、Redo 還原。每次修改都記錄是誰、何時做的。

    快速開始

    需要 Python 3.8 以上,只用標準函式庫,不用另外安裝東西。

    搭配 Claude Code(終端機、VS Code 或 JetBrains),安裝一次就好:

    /plugin marketplace add Ashur5457/duetkifu
    /plugin install duetkifu@duetkifu

    之後請 Claude 針對你的資料夾啟動 Duetkifu(/duetkifu skill)。

    不想裝外掛的話,直接執行:

    python duetkifu.py "你的資料夾路徑"

    瀏覽器會自動開啟並連到該資料夾。原始資料只會被讀取,紀錄和報告存在資料夾內的 duetkifu/ 子資料夾。介面支援英文、繁體中文、簡體中文、日文、韓文、西班牙文。

    常見問題

    Duetkifu 是什麼?給與 AI agent 一起做研究的人用的開源研究紀錄與報告審閱工具,保存每一步(含死路)以及每個數字的出處。

    支援哪些 AI agent?Claude Code、Codex、Gemini CLI、Cursor,或任何能編輯 JSON 的 agent(見 AGENTS.md)。

    要錢嗎?免費,MIT 授權。目前是 v0.8 原型,歡迎試用後回報問題。

    和實驗筆記或聊天紀錄有什麼不同?每個 move 有結構(接續哪步、理由、結果、失敗原因),數字會和重新計算、檔案指紋比對,整條路徑畫成可編輯的樹。

    試用有問題或想要的功能,歡迎到 GitHub 開 issue,或直接留言給我。

    關鍵字:研究紀錄、研究日誌、AI agent 實驗筆記、死路、負面結果、可重現性、資料來源追蹤、人機協作研究、Claude Code 外掛、Codex、Gemini CLI、棋譜、Duetkifu、Duetsheet。

    2026年9月27日 星期日

    Duetsheet: An Interactive Research Report for Human–AI Collaboration

     AI agents are becoming useful for data analysis and scientific reporting, but one problem remains:

    Humans still have a poor interface for telling AI exactly what is wrong.

    Instead of describing “the points in the upper-left of Figure 2” in a chat window, I built Duetsheet.

    Duetsheet is an open-source interactive HTML report where humans and AI agents work on the same document.

    Researchers can:

    • edit text and charts directly

    • click or lasso data points and leave comments

    • track human and AI revisions

    • trace figures back to datasets, scripts, and raw data

    Chart annotations can store the actual data range and point IDs, not just screen coordinates, so an AI agent can understand exactly which measurements the researcher is referring to.

    Reports are stored as plain JSON with a published JSON Schema and AGENTS.md, making Duetsheet compatible with file-capable agents such as Claude Code, Codex, Gemini CLI, and Cursor.

    The main use case is simple:

    AI drafts a data-heavy report, and a domain expert reviews and corrects it figure by figure.

    Duetsheet grew out of my own scientific research workflow and is still being actively tested and developed.

    GitHub:
    https://github.com/Ashur5457/duetsheet

    2026年9月24日 星期四

    Duetsheet:給研究者與 AI Agent 協作的互動式研究報告

    近年我常使用 AI agent 協助資料分析、繪圖與研究報告整理,但實際使用後,我一直遇到一個問題:

    AI 很會產生報告,但人類很難精確告訴它「哪裡錯了」。

    例如一張圖裡某幾個資料點有問題,在聊天視窗裡只能說「左上角那幾個點不太對」,但 AI 並不知道你實際指的是哪些資料。

    因此我做了 Duetsheet。

    Duetsheet 是一個開源的 interactive HTML report,讓研究者和 AI agent 直接在同一份報告上工作。

    研究者可以:

    • 直接修改文字與圖表

    • 點選或圈選資料點並留下註解

    • 查看 human / AI 每一次修改

    • 追蹤 figure、dataset、script 與 raw data 的來源關係

    對圖表的 annotation 不只保存畫面座標,也可以保存實際的 data range 與 data point IDs,因此 AI 能更準確理解研究者指出的問題。

    Duetsheet 的 report 使用 plain JSON,並提供 JSON Schema 與 AGENTS.md,因此不綁定特定模型,可與 Claude Code、Codex、Gemini CLI、Cursor 等 agent 搭配。

    它最適合的情境是:

    AI 先產生 data-heavy report,再由 domain expert 逐張圖、逐項資料檢查與修正。

    這個專案最初就是從我的 scientific research workflow 出發,目前仍持續在實際研究工作中測試與改進。

    GitHub:
    https://github.com/Ashur5457/duetsheet

    2026年4月8日 星期三

    Claude Pro 用戶的 Token 消耗實測紀錄:你的額度到底去哪裡了?

    最近 Anthropic 推出了一個限時活動,Pro 用戶可以免費領 $20 的 Extra Usage(超量使用額度),截止日是 4 月 17 日。趁這個機會,整理一下 Claude Pro 的 token 消耗概況,讓有需要的人參考。


    1. Pro 方案的基本額度架構

    Claude Pro 並不是單純按 token 數計費,而是採用「五小時 session 重置制」:

    • 每個五小時視窗有固定配額,用完就鎖住,等重置
    • 另外還有每週用量上限(seven-day cap),就算每個 session 都用滿,週上限也會擋住
    • 社群實測每個 session 大約有 44,000 tokens 的容量,換算大概是 10~40 則訊息,視對話複雜度而定

    而 Extra Usage(超量用量)是在配額用完之後,按 API 費率計費的付費補充機制,費率如下:

    模型Input(per 百萬 tokens)Output(per 百萬 tokens)
    Sonnet 4.6(主力)$3$15
    Opus 4.6(最強)$5$25
    Haiku 4.5(輕量)$1$5

    以一般對話 input:output ≈ 3:1 的比例估算,$20 大約可換 3~5 百萬 tokens。


    2. 各功能 Token 消耗估算

    以下數據來自實際使用 Cowork(桌面協作功能)和 Gmail 連接器的經驗整理,供參考。

    2-1. 一般對話(輕量)

    每次消耗:2,000 ~ 8,000 tokens

    最基本的問答,消耗最低。拿來解釋概念、腦力激盪都沒問題,一個 session 配額可以撐很久。


    2-2. Gmail 連接器 / 撈信件資料(中量)

    操作消耗估算
    搜尋幾封信(只看標題 / 摘要)5,000 ~ 15,000 tokens
    讀取完整郵件內容(每封)+3,000 ~ 10,000 tokens
    大量搜尋 + 多輪整理(如稅務信整理)50,000 ~ 120,000 tokens

    注意事項: Gmail、Google Calendar 這類 MCP 連接器,每次呼叫都會把完整的工具定義(schema)注入 context,光是「開啟連接器」這個動作本身就有固定 token 開銷,跟你有沒有真的搜尋到東西無關。

    建議:先用精準關鍵字縮小搜尋範圍,確定要哪幾封再讀全文,避免廣撒網式多輪搜尋。


    2-3. Cowork 生成 Word / Excel / PowerPoint(中高量)

    功能單次消耗估算每次來回修改
    Word(.docx)15,000 ~ 40,000+5,000 ~ 15,000
    Excel(.xlsx)10,000 ~ 30,000+5,000 ~ 10,000
    PowerPoint(.pptx)20,000 ~ 60,000+8,000 ~ 20,000

    消耗偏高的主要原因是:Cowork 每次生成文件,都需要先讀取一份很長的內建教學文件(SKILL.md),才能知道怎麼用 Python 正確生成對應格式的檔案。PowerPoint 最貴,因為教學文件最大、輸出頁面最多。

    建議:一次給清楚的需求,盡量減少來回修改次數。 每一輪修改都要把整個 context 重新送一遍,成本不低。


    2-4. Cowork 寫代碼 / 跑分析(高量)

    情境消耗估算
    簡單腳本(單一任務、單檔)10,000 ~ 30,000 tokens
    中型任務(含除錯往返 2~3 輪)30,000 ~ 100,000 tokens
    複雜研究代碼(含讀取多份文件、資料分析)50,000 ~ 150,000+ tokens

    對於有讀取文件、執行 bash 指令的任務,工具輸出的原始結果會完整進入 context,例如 cat 一份 500 行的文件,那 500 行就全部塞進去了。Session 越長、context 越肥,後面每一輪的成本就越高。


    3. 消耗彙整對照表

    功能消耗等級單次估算(tokens)
    一般問答🟢 輕量2,000 ~ 8,000
    Gmail 搜尋信件🟡 中量5,000 ~ 15,000
    Gmail 大量整理🟠 中高量50,000 ~ 120,000
    生成 Excel🟠 中高量10,000 ~ 30,000
    生成 Word🟠 中高量15,000 ~ 40,000
    生成 PowerPoint🔴 高量20,000 ~ 60,000
    寫代碼(簡單)🟠 中高量10,000 ~ 30,000
    寫代碼(複雜)🔴 高量50,000 ~ 150,000+

    4. 省 Token 的實用技巧

    ① CLAUDE.md 保持精簡
    如果你有用 Cowork 並設定了 CLAUDE.md,這份文件的內容會在每一輪對話開始時全部注入 context。文件有 5,000 tokens,等於每次互動都先扣掉 5,000 tokens,還沒看你的問題就先燒掉了。建議主文件只放路徑指引,細節拆到子文件,用到再讀。

    ② 一次把需求說清楚
    每一個 follow-up 訊息都會把整個對話歷史重新送一遍。如果有三個相關問題,合併成一則訊息問,比分三次問省很多。

    ③ MCP 工具搜尋要精準
    Gmail 等連接器的搜尋結果如果太廣,返回大量原始資料進 context 會非常貴。先縮小搜尋範圍再讀全文。

    ④ 善用離峰時段(針對日本用戶)
    Anthropic 在平日峰值時段(ET 08:00–14:00,即 JST 22:00–04:00)會限縮 session 配額。日本時間的白天(JST 04:00–22:00)和週末全天屬於離峰,配額相對寬鬆,適合安排大型任務。

    ⑤ 避免在同一 session 重複讀取同一檔案
    Cowork 在同一 session 內如果多次讀取相同文件,每次都是全額計費,這是最常見的浪費來源之一。


    5. $20 免費額度對 Pro 用戶的實際意義

    這次的 $20 免費額度是 Extra Usage 形式,不是訂閱配額的加倍,而是「配額用完後的應急緩衝」:

    • 純對話用途:大約可以多撐數百到上千則訊息
    • Cowork 生成文件為主:大約可以多撐 20~60 個工作 session
    • 重度代碼分析(每 session 燒 100K+ tokens):大約 10~30 個 session

    有效期限是領取後 90 天,過期歸零。

    領取步驟(僅限網頁版,4/17 前):

    1. 登入 claude.ai 網頁版
    2. 進入 Settings > Usage
    3. 找到 Extra usage,開啟開關
    4. 點擊「Claim」領取
    5. 記得關掉 Auto-reload(自動加值),否則額度用完系統會自動刷卡補值

    以上數據為估算值,實際消耗會因對話長度、工具使用頻率、模型選擇而有所不同。如果有不同的使用經驗歡迎留言分享。

    2025年3月9日 星期日

    Design of Electrolyte Using Deep Eutectic Solvents for High-Performance Rechargeable Nickel-Iodine Batteries

     Abstract

    Rechargeable nickel-ion batteries (RNiBs) have attracted significant attention because of their high volumetric density, low cost, environmental friendliness, and easy recyclability. In this study, a rechargeable nickel-iodine battery using a rational design of a deep eutectic solvent (DES) electrolyte based on a conversion reaction mechanism is first demonstrated. The rechargeable Ni-I2 battery with the DES electrolyte delivered a specific capacity of 201 mAh g−1 with a coulombic efficiency of 82.5% over 65 cycles at a current density of 0.3 A g−1. The energy storage mechanism can be attributed to I+/I− redox chemistry, which has been validated by ex situ Raman, X-ray photoelectron spectroscopy (XPS) and X-ray absorption spectroscopy (XAS). The study provides an avenue for exploring rechargeable nickel-ion batteries with DES electrolytes based on the conversion reaction mechanism.

     

    2025年2月17日 星期一

    深度學習模型權重檔案格式與存放目錄

     隨著深度學習模型的發展,越來越多的開發者透過 GitHub 與 Hugging Face 分享模型權重,以便其他人可以下載並加以應用。但不同的深度學習框架有各自的儲存格式與資料夾結構,因此了解這些規範能幫助我們更快速找到所需的模型。

    1. 常見的深度學習權重檔案格式

    不同的深度學習框架使用不同的檔案格式來儲存模型的權重,以下是最常見的副檔名:

    副檔名

    用途

    對應框架

    .bin

    PyTorch 模型權重 (Hugging Face)

    PyTorch

    .pth / .pt

    PyTorch 權重 (state_dict 或完整模型)

    PyTorch

    .safetensors

    更安全的 PyTorch 權重存儲格式

    PyTorch, Hugging Face

    .pb

    TensorFlow Frozen Graph

    TensorFlow

    .ckpt

    TensorFlow 或 PyTorch 的 Checkpoint

    TensorFlow, PyTorch

    .h5

    Keras/TensorFlow 權重

    TensorFlow, Keras

    .tflite

    TensorFlow Lite 模型

    TensorFlow Lite

    .msgpack

    Chainer 權重存儲格式

    Chainer

    .npz

    JAX 或 NumPy 存儲格式

    JAX, NumPy

    .onnx

    ONNX 格式,方便跨框架使用

    ONNX


    當下載 Hugging Face 或 GitHub 上的模型時,可以根據這些副檔名來判斷模型的格式並選擇合適的框架來載入。


    2. 深度學習模型權重的儲存目錄

    不同的框架與專案通常會將模型權重存放在特定的目錄中,以下是最常見的結構與對應的儲存位置:

    (1) Hugging Face (transformers, diffusers, sentence-transformers 等)

    Hugging Face 的模型通常儲存在 model 相關的資料夾下,例如:

    /model
      ├── config.json
      ├── pytorch_model.bin  # PyTorch 權重
      ├── model.safetensors  # SafeTensors 權重
      ├── tf_model.h5        # TensorFlow 權重
      ├── tokenizer.json
      ├── special_tokens_map.json
    

    有些大型模型(如 LLaMA)會有多個拆分的 .bin 權重檔案:

    /model
      ├── pytorch_model-00001-of-00003.bin
      ├── pytorch_model-00002-of-00003.bin
      ├── pytorch_model-00003-of-00003.bin
      ├── tokenizer.json
      ├── config.json
    

    📌 相關目錄: /model/, /weights/, /checkpoints/, /snapshots/


    (2) PyTorch(GitHub 上常見的專案結構)

    PyTorch 模型的權重通常儲存在 weights 或 checkpoints 目錄:

    /project_root
      ├── models/
      │   ├── model.py
      │   ├── __init__.py
      ├── weights/
      │   ├── best_model.pth
      │   ├── last_checkpoint.pth
      ├── checkpoints/
      │   ├── epoch_10.pth
      │   ├── epoch_20.pth
    

    📌 相關目錄: /weights/, /checkpoints/, /models/, /logs/


    (3) TensorFlow/Keras

    TensorFlow 和 Keras 的權重通常儲存在 checkpoints 或 saved_model 目錄:

    /project_root
      ├── checkpoints/
      │   ├── model.ckpt.index
      │   ├── model.ckpt.data-00000-of-00001
      │   ├── checkpoint
      ├── saved_model/
      │   ├── assets/
      │   ├── variables/
      │   ├── saved_model.pb
    

    📌 相關目錄: /checkpoints/, /saved_model/, /logs/


    (4) ONNX(跨框架模型)

    ONNX 模型通常存放在 onnx_models 或 exported_models 目錄:

    /project_root
      ├── onnx_models/
      │   ├── model.onnx
      ├── exported_models/
      │   ├── model.onnx
    

    📌 相關目錄: /onnx_models/, /exported_models/


    (5) 擴散模型(Stable Diffusion, ControlNet)

    擴散模型通常使用 .safetensors 或 .ckpt 格式,並存放在 models 目錄中:

    /stable-diffusion
      ├── models/
      │   ├── stable-diffusion-v1-4.ckpt
      │   ├── stable-diffusion-v2.safetensors
      ├── configs/
      │   ├── v1-inference.yaml
    

    📌 相關目錄: /models/, /diffusion_models/


    總結

    框架/類型

    常見儲存目錄

    Hugging Face

    /model/, /weights/, /checkpoints/, /snapshots/

    PyTorch

    /weights/, /checkpoints/, /models/, /logs/

    TensorFlow

    /checkpoints/, /saved_model/, /logs/

    ONNX

    /onnx_models/, /exported_models/

    擴散模型

    /models/, /diffusion_models/


    2025年1月13日 星期一

    深度學習中的稀疏性:提升效率還是削弱能力?

     

    在深度學習領域,「稀疏性(Sparsity)」是一個關鍵概念,它指的是數據或模型參數中有許多值為零的特性。這種特性可以提升計算效率、減少記憶體需求,甚至提高模型的泛化能力。但這是否意味著「零越多越好」呢?其實,關鍵在於如何適當地控制稀疏性,以達到最好的平衡。本文將介紹深度學習中幾種常見的稀疏性類型,以及它們在實際應用中的影響。

    1. 稀疏性類型與應用

    (1) 參數稀疏性(Model Sparsity)

    指的是神經網路中的權重矩陣大部分為零。這可以透過 L1 正則化(Lasso)、剪枝(Pruning) 或 低秩分解(Low-rank Factorization) 來實現。

    舉例: 假設一個神經網路的權重矩陣如下:

    W=[0.500.2000−0.300.8]W = \begin{bmatrix} 0.5 & 0 & 0.2 \\ 0 & 0 & 0 \\ -0.3 & 0 & 0.8 \end{bmatrix}

    這裡有 6 個元素為 0(總共 9 個參數),稀疏度為 66.7%。這樣的矩陣可以減少儲存需求,並透過稀疏矩陣運算提升計算速度。


    (2) 激活稀疏性(Activation Sparsity)

    當使用 ReLU(Rectified Linear Unit)激活函數時,負數輸入會變成 0,導致許多神經元「沉默」。

    舉例: 輸入矩陣 XX:

    X=[2−10.5−3041−2−0.7]X = \begin{bmatrix} 2 & -1 & 0.5 \\ -3 & 0 & 4 \\ 1 & -2 & -0.7 \end{bmatrix}

    經過 ReLU 激活後:

    ReLU(X)=[200.5004100]\text{ReLU}(X) = \begin{bmatrix} 2 & 0 & 0.5 \\ 0 & 0 & 4 \\ 1 & 0 & 0 \end{bmatrix}

    這裡產生了 4 個零值(共 9 個元素),稀疏度約為 44.4%。這有助於減少計算,但若太多神經元變為 0,可能影響模型學習能力。


    (3) 特徵稀疏性(Feature Sparsity)

    指的是輸入數據本身為稀疏的,例如 自然語言處理(NLP) 的詞袋模型(BoW)、推薦系統的用戶-物品互動矩陣等。

    舉例: 詞頻向量(Bag of Words):

    BoW=[1005002]\text{BoW} = \begin{bmatrix} 1 & 0 & 0 & 5 & 0 & 0 & 2 \end{bmatrix}

    只有 3 個非零值,表示這段文字只包含 3 個詞。這樣的稀疏特徵能夠壓縮存儲並提升計算效率。


    (4) 梯度稀疏性(Gradient Sparsity)

    在深度學習訓練中,部分權重的梯度可能接近 0,意味著它們對損失函數的貢獻很小。

    舉例: 梯度矩陣:

    Gradient=[0.010−0.0200000.030]\text{Gradient} = \begin{bmatrix} 0.01 & 0 & -0.02 \\ 0 & 0 & 0 \\ 0 & 0.03 & 0 \end{bmatrix}

    在 分散式訓練 時,僅傳輸非零梯度可減少通信成本,提高計算效率。


    (5) 注意力稀疏性(Sparse Attention)

    在 Transformer 模型(如 BERT, GPT)中,自注意力機制計算量為 O(n2)O(n^2)。透過「稀疏注意力」,模型可聚焦於關鍵資訊,減少計算量。

    舉例:

    A=[0.10.30.050.020.00.60.00.00.00.20.00.00.050.00.00.4]A = \begin{bmatrix} 0.1 & 0.3 & 0.05 & 0.02 \\ 0.0 & 0.6 & 0.0 & 0.0 \\ 0.0 & 0.2 & 0.0 & 0.0 \\ 0.05 & 0.0 & 0.0 & 0.4 \end{bmatrix}

    這樣的設計可降低 Transformer 計算複雜度,提升運算效率。


    2. 0 越多越好嗎?

    許多人會問:「如果讓更多參數變成 0,是否代表更好的模型?」答案是否定的。過度稀疏會導致 信息丟失,影響模型的表現。

    適當稀疏與過度稀疏的影響

    應用場景

    適當稀疏的好處

    過度稀疏的風險

    模型壓縮(剪枝)

    減少模型大小,加快運算

    削弱表達能力,影響準確度

    ReLU 激活

    過濾無效資訊,提高計算效率

    過多神經元變成 0,影響學習

    NLP 稀疏注意力

    只關注重要詞,提高效率

    忽略重要詞,影響理解

    推薦系統(特徵稀疏)

    加速運算,減少存儲需求

    缺少重要的交互信息


    3. 如何控制稀疏度?

    要讓模型既能利用稀疏性提升效率,又不會過度影響學習能力,可以考慮以下方法:

    1. 逐步調整剪枝比例(如 30%、50%、70%)來測試影響。
    2. 使用 L1 正則化 來鼓勵但不強制 0 值。
    3. 採用動態稀疏技術(Dynamic Sparsity),讓模型在訓練中自行選擇要稀疏的部分。

    結論

    稀疏性是一種強大的工具,能夠提升深度學習的運算效率,但「零越多越好」的想法是錯誤的。關鍵在於 適當平衡稀疏與模型表達能力,才能在效率與準確度之間取得最佳效果。

    你是否在使用稀疏性來加速你的深度學習模型?歡迎在留言區分享你的經驗!

    2024年12月8日 星期日

    [文章轉貼] PyTorch GPU加速指南:如何使用CUDA進行基本操作

    文章轉自微信公眾號:阿旭演算法與機器學習

    文章原始連結:https://mp.weixin.qq.com/s/OJ1_S0b39VIr7_4kfn5Yaw

    引言

    CUDA(Compute Unified Device Architecture)是NVIDIA專有的平行運算平台和程式設計模型。使用CUDA SDK,開發人員可以利用他們的NVIDIA GPU(圖形處理單元),從而使他們能夠在通常的程式設計工作流程中引入基於GPU的平行處理能力,而不是通常的基於CPU的順序處理能力。

    隨著近年來深度學習的興起,可以看到模型訓練中涉及的各種運算,如矩陣乘法,求逆等,可以在很大程度上並行化,以獲得更好的學習表現和更快的訓練週期。因此,許多像Pytorch這樣的深度學習函式庫使用戶能夠使用一組介面和實用程式函數來利用GPU。本文將介紹在任何包含支援CUDA的GPU的系統中設定CUDA環境,並簡要介紹使用Python的Pytorch庫中提供的各種CUDA操作。

    查看GPU支援的CUDA版本

    在cmd控制台輸入navidia-smi查看GPU支援的最高CUDA版本:




    如上圖所示,最高支援的CUDA版本為12.5,版本可以向下相容。因此安裝的CUDA版本必須小於或等於12.5版本。

    安裝GPU版Pytorch

    首先,透過官方Nvidia CUDA相容性清單檢查其係統的GPU,以確保其GPU是否啟用CUDA。 Pytorch透過提供一個很好的使用者友善介面,讓您選擇作業系統和其他要求,讓CUDA安裝過程非常簡單,如下圖所示。根據我們的計算機,我們將根據下圖中給出的規格進行安裝。

    參考Pytorch的官方連結:https://pytorch.org/get-started/locally/,根據他們的電腦規格選擇規格。我們還建議在安裝後完全重新啟動系統,以確保工具包的正常運作。




    Pytorch安裝頁面截圖

    ❝

    pip3 install torch==1.9.0+cu102 torchvision==0.10.0+cu102 torchaudio=0.9.0 -f https://download.pytorch.org/whl/torch_stable.html

    在Pytorch中開始使用CUDA

    安裝後,我們可以使用torch.cuda介面使用Pytorch與CUDA互動。我們將使用以下函數:

    ❝

    文法:

    1. torch.version.cuda() :傳回目前安裝的軟體包的CUDA版本
    2. torch.cuda.is_available():如果您的系統支援CUDA,則傳回True,否則傳回False
    3. torch.cuda.current_device():傳回目前裝置的ID
    4. torch.cuda.get_device_name(device_ID):傳回ID = 'device_ID'的CUDA裝置的名稱

    代碼:

    import torch

    print(f"CUDA是否可用? {torch.cuda.is_available()}")
    print(f"当前CUDA 版本: {torch.version.cuda}")

    # Storing ID of current CUDA device
    cuda_id = torch.cuda.current_device()
    print(f"当前CUDA ID:{torch.cuda.current_device()}")

    print(f"CUDA设备名称:{torch.cuda.get_device_name(cuda_id)}")

    輸出:



    使用CUDA處理張量

    為了透過CUDA互動Pytorch張量,我們可以使用以下實用函數:

    ❝

    文法:

    • tensor.device:傳回「Tensor」所在的裝置名稱
    • Tensor.to(device_name):傳回「device_name」指定的裝置上的「Tensor」的新實例:「cpu」表示CPU,「cuda」表示支援CUDA的GPU
    • tensor.cpu():將「Tensor」從目前裝置傳輸到CPU

    為了示範上述函數,我們將建立一個測試張量並執行以下操作:

    檢查張量的目前設備並應用張量操作(平方),將張量傳輸到GPU並應用相同的張量操作(平方),並比較2個設備的結果。

    代碼:

    import torch

    # Creating a test tensor
    x = torch.randint(1, 100, (100, 100))

    # Checking the device name:
    # Should return 'cpu' by default
    print(x.device)

    # Applying tensor operation
    res_cpu = x ** 2

    # Transferring tensor to GPU
    x = x.to(torch.device('cuda'))

    # Checking the device name:
    # Should return 'cuda:0'
    print(x.device)

    # Applying same tensor operation
    res_gpu = x ** 2

    # Checking the equality
    # of the two results
    assert torch.equal(res_cpu, res_gpu.cpu())

    輸出:

    cpu
    cuda : 0

    使用CUDA處理深度學習模型

    一個好的Pytorch實踐是產生與裝置無關的程式碼,因為某些系統可能無法存取GPU,只能依賴CPU,反之亦然。完成後,可以使用以下函數將任何機器學習模型傳輸到所選設備上

    ❝

    用法: Model.to(device_name):

    傳回:「device_name」指定的裝置上的機器學習「Model」的新實例:「cpu」表示CPU,「cuda」表示啟用CUDA的GPU

    在本例中,我們從torchvision.models實用程式匯入預先訓練的Resnet-18模型,讀者可以使用相同的步驟將模型傳輸到所選設備。

    代碼:

    import torch
    import torchvision.models as models

    # Making the code device-agnostic
    device = 'cuda' if torch.cuda.is_available() else 'cpu'

    # Instantiating a pre-trained model
    model = models.resnet18(pretrained=True)

    # Transferring the model to a CUDA enabled GPU
    model = model.to(device)