GitHub
335 articles in the TruthFoundry index. Each links out to the original.
- llama.cpp removes mp_21 from default musa architecture · Developer Xiaodong Ye removes mp_21 from default musa architectures in llama.cpp.
- llama.cpp Fixes WebGPU Crash by Adding GGML_OP_DUP Support · llama.cpp resolves a WebGPU crash by adding support for the GGML_OP_DUP operation in the mixed batch path.
- llama.cpp release b11521 sets default server port to 9931 · llama.cpp release b11521 changes default server port to 9931 via GitHub commit.
- llama.cpp project publishes software release b11518 on GitHub · ggml-org has tagged and released version b11518 of the llama.cpp software.
- llama.cpp release b11514 published with Musa FWHT fix · ggml-org tagged llama.cpp release b11514 including a Musa FWHT fix.
- llama.cpp project publishes verified release tag b11510 · ggml-org has tagged verified software release b11510 for the llama.cpp project.
- llama.cpp fixes CUDA CCCL version guard regression on major rollover · llama.cpp developer fixes a CUDA build regression caused by incorrect CCCL version comparison logic
- llama.cpp Adds Multi-GPU MoE Cache Support · llama.cpp developer am17an adds support for MoE cache over multiple GPUs.
- llama.cpp Adds Robust Checkpoint Handling for Slot Save/Restore · llama.cpp developers update server to handle context checkpoints across slot save and restore operations.
- Llama.cpp Vulkan Fix Resolves Top-K Errors for Extreme Inputs · Llama.cpp developer Gianni Cor fixes Vulkan TOP_K handling for infinite and negative values.
- llama.cpp release b11504 published with Vulkan feature update · ggml-org released llama.cpp version b11504 including Vulkan sparse FA improvements.
- bri-prism tags commit with SYCL FWHT optimizations · Developer bri-prism tagged a commit implementing FWHT optimizations for SYCL in the llama.cpp project.
- hex-mmadd optimization avoids alignment assumptions in fused bias-add · A commit updates hex-mmadd to not assume aligned read/write during fused bias-add operations.
- Llama.cpp b11499 Tagged with CUDA Roll Optimization · The bertaye team tagged commit b11499 for a CUDA optimization allowing non-contiguous roll operations.
- llama.cpp release b11498 published with CUDA kernel fix · ggml-org published llama.cpp release b11498 fixing CUDA norm family kernel limits.
- LLama.cpp Adds Fused Delta-Net Alpha Gate for SYCL · LLama.cpp developers merge a commit fusing the delta-net alpha gate operations for SYCL backend support.
- LLama.cpp Adds SYCL Bulk Upload Stage for Model Loading · LLama.cpp introduces a SYCL-based ring buffer stage to optimize bulk model uploads.
- llama.cpp release b11495 adds reranker classifier activation support · llama.cpp release b11495 implements classifier activation support for reranker models
- llama.cpp Optimizes CUDA Performance with New GDN State Columns · Developer SongXiaoXi updates llama.cpp to assign four GDN state columns per warp for improved CUDA efficiency.
- llama.cpp Adds Grouped MoE XMX GEMM Support via SYCL · llama.cpp project introduces grouped MoE XMX GEMM optimizations for SYCL backend.
- llama.cpp removes duplicate SYCL block-size defines · A developer fixed duplicate block-size definitions in SYCL operation headers for llama.cpp.
- llama.cpp release b11490 adds Hexagon buffer allocation improvements · llama.cpp release b11490 implements Hexagon buffer and tensor handling updates.
- llama.cpp Adds LiquidAI d1-omni 600M Decision Model · The llama.cpp project merged a pull request adding the LiquidAI d1-omni 600M decision model.
- Llama.cpp Adds Hexagon Support for Tiled Q4_K and Q6_K Quantization · Llama.cpp developer kurquhar adds support for tiled Q4_K and Q6_K GET_ROWS on Hexagon processors.
- Llama.cpp developer improves Hexagon GELU accuracy · Llama.cpp developer kurquhar releases a commit improving GELU accuracy for Hexagon processors.
- GitHub fixes Jinja parser bug for TranslateGemma chat model · A developer fixed a Jinja parser issue in the TranslateGemma chat model on GitHub.
- LiquidAI d1-3B decision model added to llama.cpp · Developer tdakhran adds LiquidAI d1-3B model to llama.cpp with LFM2 algorithm support.
- llama.cpp Adds CUDA FWHT Kernels for Block Widths Above 512 · llama.cpp introduces new CUDA FWHT kernels supporting block widths from 1024 to 8192 for F32 and F16 data types.
- GitHub User Adds Cohere2 Vision Support to Llama.cpp · Terrencezzj merged a commit adding Cohere2 vision model support to the Llama.cpp project.
- llama.cpp Adds GPU Cache for Host-Resident MoE Experts · llama.cpp introduces a GPU cache for Mixture of Experts models stored in host memory.