Hugging Face
61 articles in the TruthFoundry index. Each links out to the original.
- NeoMME: Efficient Multilingual Multimodal Encoder for Visual Document Retrieval · NeoMME introduces a single Transformer encoder for multilingual text and images, achieving state-of-the-art visual document retrieval.
- Fine-tuning LFM2.5-350M with GRPO improves structured output compliance · Researchers fine-tune LFM2.5-350M using GRPO to boost IFStruct benchmark scores from 22.6% to 29.7%.
- funes Adds Persistent Memory Layer to Coding Agents · funes enables coding agents to recall past reasoning across sessions and machines via a local and shareable memory layer.
- Developer Replicates Viral Watercolor AI Art Project with Open Source Tools · A developer reproduces a viral AI watercolor project using TRL and OpenEnv on Hugging Face.
- IBM and Confluent Launch Stream-Native Time Series AI Models · IBM and Confluent introduce Early Access time series foundation models for real-time forecasting and anomaly detection.
- Hugging Face releases 207 WebGPU kernels and Fleet benchmarking tool for browser AI · Hugging Face released @huggingface/kernels, a library of 207 WebGPU kernels for browser AI inference, plus Fleet, a benchmarking tool.
- Hugging Face and Voice Arena Add Hindi and Indian English to Open ASR Leaderboard · Hugging Face and Voice Arena launch Monsoon speech evaluation sets for Hindi and Indian English on the Open ASR Leaderboard.
- How to Finetune Multi-Vector Embedding Models for Domain-Specific Retrieval · A tutorial demonstrates finetuning Sentence Transformers MultiVectorEncoder models to improve retrieval on long, domain-specific documents.
- New Method Recovers Capabilities in Compressed 4-Bit LLMs · Researchers introduce Quantization-Aware Healing to recover capabilities in compressed large language models.
- https://huggingface.co/blog/gradio-workflow-guide
- New tests expose speech recognition models optimizing to benchmarks, not real-world audio · Qualcomm AI Research found open-source ASR models reproduce benchmark errors, overstating real-world accuracy.
- Hugging Face details hybrid search for Papers with Code using Jobs, Buckets, Inference Endpoints · Papers with Code search uses hybrid retrieval combining PostgreSQL full-text and pgvector embeddings.
- Liquid AI releases DSpark draft models for LFM2.5 family with up to 3.2x faster inference · Liquid AI released DSpark draft model checkpoints for three LFM2.5 models, claiming up to 3.18x GPU and 2.87x on-device inference speedups.
- ALTK-Evolve study: agentic memory is a per-model dose, not a switch · ALTK-Evolve's eight-model study found agentic memory works best when dosed per model: strong models get full guidelines, weaker ones get curated retrieval.
- Sentence Transformers v6.0 Adds Multi-Vector Late-Interaction Embedding Models · Sentence Transformers v6.0 adds MultiVectorEncoder for ColBERT-style late-interaction retrieval, supporting PyLate, ColBERT, and ColPali checkpoints.
- Constraint-aware GPU allocator raises utilization 33 points over FIFO scheduling · A constraint-aware GPU allocator beat a FIFO scheduler by up to 33 utilization points and 105% in priority-weighted output on identical hardware.
- State of Open AI Models 2026: Chinese Labs Dominate Frontier, Qwen as Base Model · Report: Chinese labs dominate frontier open models; Qwen emerges as base model; AMD/NVIDIA lead U.S. releases.
- AWS Strands Robots enables efficient record-train-deploy loop using Hugging Face Storage Buckets · AWS Strands Robots SDK records demos, streams from Hugging Face Storage Buckets, trains, and deploys to hardware.
- Open Reproduction Hackathon Examines 2,226 ICML Papers, Finds Falsifications · The ICML 2026 Open Reproductions hackathon reused coding agents to verify 2,226 papers, finding 23% had falsified claims.
- AI2's OlmoEarth Studio adds custom embedding exports for Earth observation analysis · OlmoEarth Studio now exports custom embedding vectors for similarity search, segmentation, and change detection.
- ALTK-Evolve matches ACE agent accuracy at a fraction of inference tokens · ALTK-Evolve reports same-or-better accuracy than ACE on AppWorld while using 40-85% fewer tokens per task.
- NVIDIA Magpie TTS open-weights model expands to 12 languages with 32ms first-audio latency · NVIDIA Magpie TTS, a 364M-parameter open-weights model, adds Arabic, Korean, and Brazilian Portuguese, achieving 32ms first-audio latency on B200.
- CompactifAI paper cuts LLM distillation VRAM with offline top-K logits and fused chunked KL loss · CompactifAI's new distillation method caches teacher top-K logits and fuses a chunked KL loss, cutting memory up to 15.6x on single GPUs.
- Meta launches Muse Glimmer, open-source 30B multimodal model for agentic AI · Meta released Muse Glimmer, a 30B-parameter Apache-2.0 multimodal model for local agentic tasks, with Hugging Face day-0 support.
- Baseten joins Hugging Face Inference Providers for serverless LLM access · Hugging Face adds Baseten as a supported Inference Provider, enabling serverless access to LLMs like DeepSeek V4 Flash.
- GPU Utilization Becomes the New Binding Constraint in Enterprise AI Economics · Enterprise AI profitability increasingly depends on keeping GPUs utilized, mirroring airlines' aircraft utilization problem.
- NVIDIA Launches Cosmos-H-Dreams, Real-Time Generative Simulator for Surgical Robotics · NVIDIA's Cosmos-H-Dreams generates real-time surgical simulation at ~160 fps on a single RTX PRO 6000 GPU.
- Hugging Face Diffusers Adds Native Loading for Nunchaku 4-bit Diffusion Models · Diffusers now loads Nunchaku Lite 4-bit quantized diffusion transformers via from_pretrained, cutting VRAM to about 12 GB on RTX 5090.
- Pollen Robotics launches Grabette, open handheld robot data recorder · Pollen Robotics releases Grabette, an open-source handheld gripper that records robot manipulation demonstrations without a robot.
- DharmaOCR Specialized for Brazilian Portuguese Outperforms Newer OCR Models on Benchmark · DharmaOCR, specialized for Brazilian Portuguese, scored 0.925 on a Portuguese-only benchmark, beating Mistral OCR4 at 0.798 and Unlimited-OCR at 0.7587.