Snorkel AI
17 articles in the TruthFoundry index. Each links out to the original.
- Continual Learning Bench Measures AI System Improvement Through Experience · Parth Asawa introduces Continual Learning Bench to evaluate AI systems' improvement through sequential experience
- New T² Scaling Laws Show Overtrained Small Models Optimal for Reasoning · Research indicates models optimized for test-time reasoning should be heavily overtrained rather than scaled up.
- Milestone-Based Evaluation Preserves Progress in Long-Horizon AI Agents · New benchmarks use milestone-based evaluation to capture partial progress in complex AI agent workflows.
- Snorkel AI Builds Realistic Enterprise Environments to Train Agents · Snorkel AI develops simulated enterprise environments to train agents on complex, domain-specific workflows.
- Claude Opus 5 Leads Debugging Tasks but Fails on Reasoning · Claude Opus 5 excels at debugging but fails due to reasoning errors, not tooling issues.
- Senior SWE-Bench Reveals Coding Agents Fail 70% of Senior-Level Tasks · New benchmark shows top coding agents fail over 70% of realistic senior engineering tasks due to quality gaps.
- Grok 4.5 Outperforms GPT 5.5 and Claude Opus 4.8 on Professional Work Benchmarks · SpaceXAI's Grok 4.5 model achieves the highest pass rates on Snorkel's GDPval+ professional reasoning dataset.
- New AI Benchmark Reveals Agents Fail Real-World Economic Tasks · Researchers launch Agents' Last Exam benchmark showing AI agents fail to automate complex professional workflows.
- New Benchmark Tests AI Agents' Ability to Learn Across Tasks · Researchers launch Continual Learning Bench to evaluate if AI agents retain experience across sequential tasks.
- AI Coding Agents Re-read Entire Codebase Every Session · Researchers warn AI coding agents inefficiently re-scan codebases each session instead of learning dependencies.
- Snorkel CEO: Better Data and Benchmarks Needed to Evaluate Agentic AI · Snorkel AI CEO Alex Ratner argues that better data and benchmarks are critical to closing the evaluation gap in agentic AI.
- Comparative Judgment Outperforms Rubrics in Legal Quality Assessment · Study finds comparative judgment recovers quality ordering nearly perfectly while rubrics barely beat chance in legal tasks.
- AI Engineer London: Building Benchmarks That Shape the Field · Vincent Sunn Chen discusses the science and art of creating AI benchmarks that drive field progress.
- AI Agents Fail at Building Schematics in KiCad Benchmark · A new benchmark reveals frontier AI agents succeed only in editing existing schematics but fail at building them from scratch.
- Federal Leaders Demand Trust Evidence for AI Deployments · Government and security leaders argue AI pilots fail due to lack of reliability evidence, not accuracy.
- Stanford Researchers Propose 'Collaborative Gym' Framework for Human-Agent Work · Yijia Shao presents a new framework enabling bidirectional human-agent collaboration to preserve human agency.
- ProgramBench Benchmark Reveals Frontier AI Struggles to Build Software from Scratch · John Yang's ProgramBench benchmark shows frontier models scored 0% initially at rebuilding real software without internet access.