LM: Difference between revisions

From Fundamental Ramen
Jump to navigation Jump to search
Line 48: Line 48:
     T3[Install nvtop]
     T3[Install nvtop]
     T1 --> T2
     T1 --> T2
    B1 --> T1
    B1 --> T3
   end
   end
</quickmmd>
</quickmmd>

Revision as of 03:08, 13 August 2026

Decision Tree for tinkering with Intel Arc Pro B70

File Path/images/quickmmd/LM-tinkering-b70.svg
File Size29.88 KB
Last modified20:55:10
Time of SVG loading2.23 second(s)
Time of SVG convertion0.00 second(s)
Cache DetectedNo
MD5 of incoming syntaxa606269225c62126494287925359b695
MD5 of cached syntaxb717d64ff51dccab5da53a61c94c6d48
SVG EngineKroki API
Kroki API URLhttp://kroki:8000
Mermaid syntax Referencehttps://mermaid.js.org/intro/syntax-reference.html
QuickMMD Version1.0.0 - About QuickMMD
flowchart LR
  subgraph zero
    A[Got B70]
  end

  subgraph os
    B1[Install Ubuntu 24.04]
    B2[Install Ubuntu 26.04]
    A --> B1
    A --> B2
  end

  subgraph drvier
    C1[Install oneAPI]
    C2[Install OpenVINO]
    C3[Install oneAPI]
    C4[Install OpenVINO]
    B1 --> C1
    B1 --> C2
    B2 --> C3
    B2 --> C4
  end

  subgraph runtime
    D1[Build llama.cpp for SYCL]
    D2[Build llama.cpp for OpenVINO]
    D3[Build SGLang for SYCL]
    D4[Build SGLang for OpenVINO]
    C1 --> D1
    C2 --> D2
    C1 --> D3
    C2 --> D4
  end

  subgraph model test
    E1[TODO]
    E2[TODO]
    D1 --> E1
    D2 --> E2
  end

  subgraph tools
    T1[Install pipx]
    T2[Install hf]
    T3[Install nvtop]
    T1 --> T2
    B1 --> T1
    B1 --> T3
  end

Build environment

hf (model management)

Purpose Command
cache management
hf cache list
hf cache rm <model id>
hf cache prune
fix WiFi problem
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF" --include "*UD-Q4_K_XL*"
search
# search by population
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort downloads --no-truncate --limit 25
# search for latest publish
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort created_at --no-truncate --limit 25
# search for latest tunning
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort last_modified --no-truncate --limit 25
optimize Qwen3-Coder
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 1 --format json | jq .
hf models ls -h "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF"
hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
optimize gemma-4-E4B
hf models ls -h unsloth/gemma-4-E4B-it-qat-GGUF
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mtp-gemma-4-E4B-it.gguf"
optimize gemma-4-12B
hf models ls -h unsloth/gemma-4-12B-it-qat-GGUF
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mtp-gemma-4-12B-it.gguf"
optimize Muse Glimmer
hf models ls -h unsloth/Muse-Glimmer-30B-GGUF
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q5_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q6_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "dflash-kquant.gguf"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "mmproj-kquant.gguf"
tree ~/.cache/huggingface/hub/models--unsloth--Devstral-Small-2-24B-Instruct-2512-GGUF/snapshots
tree ~/.cache/huggingface/hub/models--unsloth--Qwen3-Coder-30B-A3B-Instruct-GGUF/snapshots

Ubuntu

Agents

Wishlist