Snorkel AI
24 articles in the TruthFoundry index. Each links out to the original.
- MedPAIR Dataset Measures Physician And AI Alignment In Medical QA · Researchers present MedPAIR, a dataset comparing physician and LLM clinical reasoning relevance.
- Designing Robust RL Environments for LLM Agents · RL environments for LLM agents require reliable state, verifiable rewards, and domain-specific controls to enable effective training.
- Opus 5.5, Opus 5, Fable 5.1 AI coding benchmark results compared · Standardized benchmark tests evaluate coding performance across three frontier AI models.
- Opus 5.5, Opus 5, Fable 5.1 AI Coding Benchmark Results Published · Benchmark compares coding performance of Fable 5.1, Opus 5, and new Opus 5.5 AI models.
- Snorkel Secures $350M Series E to Build 'Data 2.0' AI Research Lab · Snorkel raises $350M Series E to advance human-AI collaboration for high-quality AI training data.
- Snorkel raises $350M Series E at $3.5B valuation for AI data platform · Snorkel raised $350M Series E led by Insight and S32, and reported $375M annualized revenue run rate.
- Grok 4.7 ranks sixth on Senior SWE-Bench with lower cost per trial · Grok 4.7 achieves a 40.0% tasteful pass@3 on Senior SWE-Bench at $0.24 per trial, ranking sixth overall.
- Snorkel AI proposes curriculum learning and introduces Terminal-Bench+ dataset for coding agents · Snorkel AI's Terminal-Bench+ provides a measured-difficulty, 50,000-task curriculum for training coding agents.
- Curriculum Learning Improves AI Coding Agent Model Training Outcomes · Progressive curriculum learning addresses skill gaps in AI coding agent development.
- OSWorld 2.0 Benchmark Shows AI Agents Fail 80% of Complex Computer Tasks · New OSWorld 2.0 benchmark reveals frontier AI agents complete only 20.6% of complex, multi-step computer tasks.
- OSWorld 2.0 Benchmark Reveals AI Agents Fail 80% of Complex Computer Tasks · OSWorld 2.0 benchmark shows frontier AI agents succeed in only 20.6% of complex, multi-step computer-use tasks.
- Fable 5.1 Outperforms Opus 5 on Debugging and Games via Efficiency · Fable 5.1 is faster and more token-efficient than Opus 5 but less reliable on build tasks.
- Fable 5.1 Outperforms Opus 5 in Efficiency but Lags in Reliability · Fable 5.1 is 36% faster and uses 58% fewer tokens than Opus 5 but fails more often on convention-heavy tasks.
- Terminal-Bench 4.0 Introduces Continuous QA for Frontier Model Evaluation · Terminal-Bench 4.0 implements a rigorous continuous QA pipeline to maintain benchmark relevance against accelerating frontier models.
- Terminal-Bench 4.0 introduces continuous QA to maintain benchmark value · Terminal-Bench 4.0 implements a continuous QA pipeline to ensure benchmark tasks remain robust and solvable.
- Terminal-Bench 3.0 Reveals Frontier AI Agents Fail Real Engineering Tasks · Terminal-Bench 3.0 shows top AI agents fail at complex streaming and ML monitoring debugging tasks.
- Terminal-Bench 3.0 Reveals Frontier AI Agents Fail Complex Engineering Tasks · Terminal-Bench 3.0 benchmark shows top AI agents struggle with complex streaming and ML monitoring debugging tasks.
- New Continual Learning Bench Aims to Measure Whether LLMs Actually Improve with Experience · Berkeley PhD student Parth Asawa presented CL-Bench, a benchmark with six domains designed to test whether LLMs improve with experience.
- Continual Learning Bench Measures AI System Improvement Through Experience · Parth Asawa introduces Continual Learning Bench to evaluate AI systems' improvement through sequential experience
- New Scaling Laws Show Overtrained Small Models Optimal for Reasoning · Train-to-Test scaling laws indicate overtrained small models are compute-optimal for reasoning tasks.
- Milestone-Based Evaluation Preserves Progress in Long-Horizon AI Agents · Researchers propose milestone-based evaluation to track partial progress in complex AI agent workflows.
- Snorkel AI Builds Realistic Enterprise Environments for Training Agents · Snorkel AI develops simulated enterprise environments to train agents on complex, domain-specific workflows.
- Claude Opus 5 Leads Senior SWE-Bench in Debugging and Bug Investigation · Claude Opus 5 ranks second on Senior SWE-Bench, excelling in debugging and bug investigation while using fewer tokens than predecessors.
- Snorkel AI Launches Senior SWE-Bench Coding Agent Benchmark · Snorkel AI released Senior SWE-Bench to test coding agents on senior engineering work.