- MXFP4/MXFP8量子化 — ローカルLLM実行の新常識
MXFP4/MXFP8はLLM向け新4ビット量子化標準。120Bパラメータモデルが1枚の80GB GPUで動作、llama.cpp/vLLMが対応。
3 minja - MXFP4/MXFP8 Quantization — Running LLMs Locally
MXFP4/MXFP8 is the new 4-bit quantization standard for LLMs. Enables 120B-param models on a single 80GB GPU via llama.cpp and vLLM.
3 minen - 2026年の主要ニューラルプロセッシングユニット
Apple ANE、Intel NPU、AMD XDNA、Qualcomm Hexagon、Google TPU、Samsung Exynos、MediaTek APU — 主要NPUを総まとめ。
3 minja - Run Shell Commands in JavaScript with Bun Shell
Learn how to run cross-platform shell commands directly in JavaScript and TypeScript using built-in Bun Shell.
1 minen - Web Security Headers — Bảo Vệ Website Từ HTTP Headers
Các HTTP security headers quan trọng: CSP, HSTS, X-Frame-Options, Permissions-Policy — cấu hình và best practices.
4 minvi - Training on Apple Neural Engine with Private APIs
ANE project reverse-engineers Apple Neural Engine private APIs to train transformers directly on ANE — no CoreML, no GPU, no Metal.
3 minen - Training trên Apple Neural Engine qua private API
Dự án ANE reverse-engineer private API Apple Neural Engine cho phép training transformer trực tiếp trên ANE — không CoreML, không GPU, không Metal.
3 minvi - vLLM — High-Performance LLM Serving Engine
vLLM is the leading open-source LLM inference engine with PagedAttention, continuous batching, FP4 quantization — up to 24x faster than naive transformers.
2 minen
Back
Blog
Page 75 - Showing 8 of 724 postsView all posts by years →