LM

From Fundamental Ramen
Revision as of 09:02, 13 August 2026 by Tacoball (talk | contribs) (→Topics)
Jump to navigation Jump to search

Topics

Agent Model Runner Driver
  • Management
  • SGLang
  • llama.cpp for SYCL
  • llama.cpp for OpenVINO
  • xe Driver
  • oneAPI
  • OpenVINO

News

Decision Tree for tinkering with Intel Arc Pro B70

File Path/images/quickmmd/LM-tinkering-b70.svg
File Size47.97 KB
Last modified21:54:49
Time of SVG loading2.31 second(s)
Time of SVG convertion0.00 second(s)
Cache DetectedNo
MD5 of incoming syntax6c1827bb6f6a512a8e4ffcb71e384f36
MD5 of cached syntax47d19aaedfbb9b5eb98f18ddfd23af9d
SVG EngineKroki API
Kroki API URLhttp://kroki:8000
Mermaid syntax Referencehttps://mermaid.js.org/intro/syntax-reference.html
QuickMMD Version1.0.0 - About QuickMMD
flowchart LR
  subgraph hardware
    A[Got B70]
  end

  subgraph os
    B1[Install<br>Ubuntu 24.04]:::node-okay
    B2[Install<br>Ubuntu 26.04]:::node-warn
    A --> B1
    A --> B2
  end

  subgraph drvier
    C1[Install<br>PPA driver]:::node-okay
    C2[Install<br>oneAPI]:::node-okay
    C3[Install<br>OpenVINO]:::node-okay
    C4[Install<br>PPA driver]:::node-warn
    C5[Cannot install<br>oneAPI]:::node-fail
    C6[Cannot install<br>OpenVINO]:::node-fail
    C1 --> C2 & C3
    C4 --> C5 & C6
    B1 --> C1
    B2 --> C4
  end

  subgraph runtime
    D1[Build llama.cpp for<br>SYCL]:::node-okay
    D2[Build llama.cpp for<br>OpenVINO]:::node-okay
    D3[Build SGLang for<br>SYCL]:::node-todo
    D4[Build SGLang for<br>OpenVINO]:::node-todo
    C2 --> D1
    C3 --> D2
    C2 --> D3
    C3 --> D4
  end

  subgraph tools
    T1[Install<br>pipx]
    T2[Install<br>hf]
    T3[Install<br>nvtop]
    T4[Install<br>screen]
    T5[Install<br>llama-swap]
    T1 --> T2
    B1 --> T1
    B1 --> T3
    B1 --> T4
    B1 --> T5
  end

  subgraph models
    M1[Muse-Glimmer-30B]
    M2[Devstral-Small-2]
    M3[Qwen3-Coder-30B]
    M4[Gemma-4-E4B]
    T2 --> M1 & M2 & M3 & M4
  end

  subgraph env
    V1[Optimize<br>.bashrc]
    V2[Optimize<br>llama-swap.yaml]
    C2 --> V1
    C3 --> V1
    T2 --> V1
    T4 ---> V1
    T5 ---> V2
  end

  subgraph service
    Z[service]
    M1 --> Z
    D1 --> Z
    V2 --> Z
  end

  classDef node-okay fill:#efe,stroke:#393
  classDef node-warn fill:#fff0e0,stroke:#d63
  classDef node-fail fill:#fee,stroke:#d33
  classDef node-todo fill:#eee,stroke:#777

Build environment

hf (model management)

Purpose Command
cache management
hf cache list
hf cache rm <model id>
hf cache prune
fix WiFi problem
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF" --include "*UD-Q4_K_XL*"
search
# search by population
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort downloads --no-truncate --limit 25
# search for latest publish
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort created_at --no-truncate --limit 25
# search for latest tunning
hf models ls --search "Muse-Glimmer" --apps llama.cpp --expand "downloads,likes,createdAt,lastModified" --sort last_modified --no-truncate --limit 25
optimize Qwen3-Coder
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 1 --format json | jq .
hf models ls -h "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF"
hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
optimize gemma-4-E4B
hf models ls -h unsloth/gemma-4-E4B-it-qat-GGUF
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mtp-gemma-4-E4B-it.gguf"
optimize gemma-4-12B
hf models ls -h unsloth/gemma-4-12B-it-qat-GGUF
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mtp-gemma-4-12B-it.gguf"
optimize Muse Glimmer
hf models ls -h unsloth/Muse-Glimmer-30B-GGUF
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q5_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q6_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "dflash-kquant.gguf"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "mmproj-kquant.gguf"
tree ~/.cache/huggingface/hub/models--unsloth--Devstral-Small-2-24B-Instruct-2512-GGUF/snapshots
tree ~/.cache/huggingface/hub/models--unsloth--Qwen3-Coder-30B-A3B-Instruct-GGUF/snapshots

Ubuntu

Agents

Wishlist