跳到主要內容

發表文章

目前顯示的是有「筆記」標籤的文章

16GB ROCm 本地 LLM-as-a-Verifier 實測數據:Qwen3.5-9B 跑起來到底多快

TL;DR: 用 Qwen3.5-9B(Q4_K)在單張 16GB AMD GPU 上跑本地 LLM-as-a-Verifier,每次 code comparison 約 0.5–2 秒,cache hit rate 66%,正確程式碼穩定給 1.0 分,有 bug 的給 0.0–0.14 分。16GB 跑 9B 模型綽綽有餘,還有空間跑其他服務。 這篇是實測續篇,前兩篇分別講了 架構設計 和 Docker 封裝 。前兩篇講的是「做什麼」跟「怎麼做」,這篇來量「到底多快」。 1. 測試環境 在一台 homelab Linux(CachyOS)上跑的: 項目 內容 GPU AMD 16GB(15.9 GiB)via ROCm 模型 Qwen3.5-9B(8.95B 參數,Q4_K ~5.7GB GGUF) 後端 llama.cpp(llama-server,ROCm HIP build) Verifier Docker container(llm-verifier),port 8010 模型 VRAM ~10.2 GB(模型本身 + 動態 KV cache) GPU 使用率 78% 系統 RAM 62 GB;使用 27 GB(swap 用了 9.3 GB) Context window 131,072 tokens MIN_SCORE 0.8 後端跑在另一台機器上,走 local Gigabit 網路連線。 2. Verifier 在幹嘛 docker-llm-as-a-verifier 包裝了 LLM-as-a-Verifier 這個研究套件,讀取 token-level log probability 來算連續分數,不是簡單的 yes/no 判斷。 Container 開了 7 個 HTTP endpoint: Endpoint Method 用途 /health GET 健康檢查 /v1/compare POST 兩組答案比對評分 /v1/select POST Best-of-N 選最佳 /v1/track POST Agent 軌跡分數追蹤 /v1/directed POST 導...

關於幸福

論語《述而》 飯疏食飲水,曲肱而枕之,樂亦在其中矣。不義而富且貴,於我如浮雲。 --- Francois de La Rochefoucauld We are more interested in making others believe we are happy than in trying to be happy ourselves. Before we set our hearts too much upon anything, let us examine how happy they are, who already possess it. --- Immanuel Kant Rules for Happiness: something to do, someone to love, something to hope for. --- John Lennon When I was 5 years old, my mother always told me that happiness was the key to life. When I went to school, they asked me what I wanted to be when I grew up. I wrote down ‘happy’. They told me I didn’t understand the assignment, and I told them they didn’t understand life. --- 妙法蓮華經卷第一 方便品第二 舍利弗。現在十方無量百千萬億佛土中諸佛世尊。多所饒益安樂眾生。是諸佛亦以無量無數方便。種種因緣譬喻言辭。而為眾生演說諸法。是法皆為一佛乘故。是諸眾生從佛聞法。究竟皆得一切種智。 --- Buddha Thousands of candles can be lighted from a single candle, and the life of the candle will not be shortened. Happiness never decreases by being shared. 《佛說四十二章經》第十章 喜施獲福 佛言:睹人施道,助之...