arXiv.org
4877 articles in the TruthFoundry index. Each links out to the original.
- Editable Visual Design System Combines Coding Agents and VLMs · New system merges VLM creativity with code-based generation for editable visual assets.
- IDRBench Benchmarks Large Language Models on Interdisciplinary Research · Researchers introduce IDRBench to evaluate state-of-the-art LLMs on interdisciplinary research tasks.
- Researcher Withdraws Causal-Counterfactual RAG Paper Due to Required Revisions · Abhay Gupta withdraws his arXiv paper on Causal-Counterfactual RAG after discovering the need for substantial methodological revisions.
- LLMs Approach Supervised Arabic Parsing with In-Context Learning · Strongest LLMs approach supervised Arabic morphosyntactic systems using retrieval-based in-context learning.
- Domain Tutorials Boost CAE Simulation Agent Performance · Adding solver tutorials significantly improves CAE simulation agent accuracy over generic harnesses.
- Intent-Aware Prompting Detects Mental Manipulation in Conversations · Researchers propose Intent-Aware Prompting to improve LLM detection of covert mental manipulation tactics using intent analysis.
- Explicit Text Imagination Outperforms Latent Visual Reasoning · Researchers found explicit text-based imagination outperforms complex latent-space reasoning in visual tasks.
- New Encoding Probe Reconstructs Language Model Representations · Researchers introduce an Encoding Probe to reconstruct internal model representations using interpretable features.
- Researchers Introduce MINT-Safe Dataset and TAD-Align Framework for MLLM Safety · New dataset and framework improve safety in multi-modal multi-turn large language model interactions.
- Research Finds Language Model Circuits Lack Task Specificity · Study reveals component-level circuits in language models are consistent but not specific to individual tasks.
- Researchers Benchmark LLMs on Urdu Idioms with New Dataset · A new benchmark evaluates frontier LLMs on translating and detecting Urdu idioms across native and Romanized scripts.
- New Benchmark CSM-MTBench Evaluates MT on Chinese Social Media · Researchers introduce CSM-MTBench to address data scarcity and metric limitations in machine translation for Chinese social media.
- Researchers formalize Dice Roll Method for auditing LLM brand recommendations · A new protocol standardizes repeated-query auditing of LLM brand recommendations using statistical metrics.
- Researchers Introduce KIDBench to Benchmark Child Safety in AI Models · New benchmark KIDBench evaluates large language models on child safety using developmental psychology.
- EDIT Framework Trains Rubric-Faithful LLM Graders via Evidence Diagnosis · Researchers propose EDIT, a two-phase framework to improve rubric adherence in LLM grading.
- RealCADBench introduces 12,632-task benchmark for parametric CAD modeling from industrial design intents · JoyIndustrial's VisCAD team presented RealCADBench, a benchmark of 12,632 tasks evaluating parametric CAD modeling via executability, IoUs, and visual-semantic identity.
- Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards · Jina-OCR-v1 improves document parsing efficiency with speculative decoding and dense verifiable rewards, achieving high throughput on low-budget GPUs.
- Study finds BioCLIP2 identifies Bangladeshi fish with 72% accuracy, Bengali prompts near chance · BioCLIP2 zero-shot model recognizes Bangladeshi freshwater fish with 72.36% accuracy on BFF-15, but Bengali prompts perform near chance.
- Multi-Perspective Adjudication Reveals Limits of Medical Hallucination Benchmarks · A study finds single-pass benchmarks undercount medical hallucinations due to annotator errors and adjudicator disagreement.
- DirBucket: Embedding-Space Watermarks Audit Third-Party RAG Reuse · DirBucket uses semantic watermarking to audit document reuse in third-party RAG systems.
- Study Reveals Frame Selection is Key Lever for Long-Video MLLMs · Research shows query-selected frames outperform uniform sampling in long-video models.
- TRACE Framework Boosts LLM Agent Consistency via Skill Bank Evolution · TRACE improves LLM agent reliability by evolving a modular skill bank without modifying model weights.
- Researchers Identify Physics Signatures in Open-Weight Language Models · Study finds open-weight LLMs encode materials-science mechanisms via specific state transformations rather than just vocabulary.
- K-Bench 01: No AI Model Clears Scientific Accuracy Threshold · K-Bench 01 reveals no frontier AI model consistently meets scientific accuracy standards on real-world agent tasks.
- New Method Improves LLM Reasoning via Skill-Conditioned Self-Distillation · Researchers propose Skill-Conditioned Gated Self-Distillation to enhance LLM reasoning using experience-derived skills.
- Activation-Keyed Momentum Optimizer Reduces Training Steps · Activation-Keyed Momentum (AK-Momentum) adapts forgetting rates to input frequency, reducing training steps for large language models.
- ParaBridge Method Improves Speech Language Model Paralinguistic Dialogue · ParaBridge uses self-distillation to enable speech language models to respond to paralinguistic cues in dialogue.
- Subspace Selection Improves LLM Probe Generalization for Deception Detection · Projecting LLM activations onto a subset of principal components enables cross-domain transfer in deception detection.
- EASE Augments Multimodal RLVR with Visual-Evidence Process Supervision · EASE improves vision-language models by aligning visual attention with annotated evidence regions during reinforcement learning.
- https://arxiv.org/abs/2609.02902