arXiv.org
4590 articles in the TruthFoundry index. Each links out to the original.
- New Symmetry Metric Fails to Predict Humor in LLMs · Researchers found a new symmetry metric failed to predict humor ratings in large language models.
- ForeDreamer: Self-Evolving Dual-Agent Architecture for Future Event Prediction · Researchers introduce ForeDreamer, a dual-agent framework that structures web evidence to improve future event prediction.
- AsmEvo Optimizes AMD GPU Kernels via Assembly-Level Agentic Verification · AsmEvo achieves significant speedups on AMD GPUs by optimizing compiled binaries through agentic assembly-level editing and verification.
- Ontology-Driven Framework Reduces Noise in Document-Level Relation Extraction · Researchers introduce an ontology-driven framework to enforce structural consistency in DocRE datasets, improving model generalization.
- TH-GNN Detects LLM-Agent Shilling Attacks via Temporal Graph Analysis · Researchers propose TH-GNN, a heterogeneous temporal graph neural network, to detect LLM-generated shilling attacks in recommender systems.
- Study Reveals Hidden Occupational Bias in Language Model Representations · Researchers found demographic attributes influence language models' internal representations of competence even when behavioral outputs appear unbiased.
- OneModel Internalizes Business Logic to Slash Latency by 50% · Researchers propose OneModel to replace modular AI pipelines with internalized knowledge, reducing latency by 50%.
- Structured Persona Extraction Improves LLM Digital Twin Accuracy · Researchers propose automatic structure discovery to enhance predictive accuracy in LLM-based digital twins.
- New Framework Reduces LLM Prompt Variance by 40% via Lexical Sensitivity Analysis · Researchers introduce an automated agent that reduces LLM performance variance by 40.7% through systematic prompt refinement.
- SSR Accelerates LLM Reasoning Traces via Self-Speculative Decoding · SSR is a training-free method that accelerates LLM reasoning traces by leveraging chain-of-thought overlap for speculative decoding.
- Self-Supervised Speech Models Track Language Convergence in Deaf Children · Researchers use speech embeddings to measure how deaf children's speech converges to adult patterns without transcription.
- Neural Networks and SVMs Classify Research Paper Quality via Text · A new benchmark classifies research papers as high-quality or flawed using only title and abstract text features.
- LLMs Struggle with Ambiguous Clinical Registry Abstraction · A study shows LLMs achieve 91.5% accuracy on clinical registry tasks but accuracy drops significantly as clinical ambiguity increases.
- Category Theory Improves Automated Research Idea Generation · Researchers propose a category theory-based algorithm to filter and validate cross-domain research analogies using paper knowledge graphs.
- ExpertIVS Framework Achieves 90% Fidelity in LLM Value Simulation · New ExpertIVS framework uses sociological agents to accurately simulate individual human value systems in LLMs.
- Therapy Bots Fail to Understand Gen Alpha Mental Health Language · AI models show a 10-14 percentage point gap in understanding youth mental health vocabulary compared to human therapists.
- Researchers Release Synthetic Bengali Speech Dataset for Telecom Customer Care · A new synthetic Bengali speech dataset for telecom customer care is released on Hugging Face with strong ASR evaluation results.
- ASTAR Framework Automates Radiology Reporting Template Creation · New AI framework ASTAR automates radiology template generation, outperforming expert-curated versions on fetal brain MRI data.
- New Framework Enables Topic Modeling of Sri Lankan Trilingual Parliamentary Debates · Researchers present an LLM-based pipeline to analyze 19,553 Sinhala, Tamil, and English speeches from Sri Lankan parliamentary debates.
- TriPLU Improves Tiny Language Models via Direct Trilinear Product FFNs · TriPLU replaces gated FFNs with direct trilinear product layers to improve tiny language model performance in low-compute regimes.
- ImmigrationReason Dataset Released for U.S. Immigration Appeals Analysis · Researchers introduce ImmigrationReason, a dataset of 12,375 U.S. immigration appeals to study legal reasoning and adjudication errors.
- EditPPT: Multi-Agent Framework Automates Faithful PowerPoint Slide Editing · EditPPT uses a multi-agent framework to automate slide editing with high accuracy and preservation fidelity.
- Intent Engine Translates Natural Language to Validated SLO Artifacts · Intent Engine architecture reduces orchestration misconfigurations by translating natural language intents into validated Service-level Objectives.
- Zero-shot LLMs Match Fine-Tuned Models Only in Specific Intent Detection Regimes · Research finds zero-shot LLMs match fine-tuned models only for out-of-scope detection, noisy inputs, or dynamic schemas.
- VA-DPO Method Enables Precise Continuous Emotion Control in Language Models · Researchers introduce VA-DPO to train language models to generate text matching specific continuous valence and arousal targets.
- LingShu: A Large-Scale Symptom-Centric Knowledge Graph Bridges TCM and Biomedicine · Researchers present LingShu, a large-scale knowledge graph linking Traditional Chinese Medicine and modern biomedicine via symptom patterns.
- Ansari: A Retrieval-Grounded Islamic AI Assistant with 140,000 Conversations · Researchers present Ansari, an Islamic AI assistant handling 140,000 queries by grounding answers in authenticated religious texts.
- JuryProbe Diagnoses Consensus Risk in Reference-Free LLM Judge Panels · JuryProbe detects correlated false negatives in LLM judge panels and routes high-risk cases to grounded verification.
- SAC-Copula Preserves Quality in Diffusion Language Model Watermarking · Researchers propose SAC-Copula, a new watermarking method for diffusion language models that preserves generation quality.
- Critical Review of LLMs in Hadith Computational Science · Scholars identify methodological gaps and data limitations in applying large language models to Hadith studies.