thenewstack.io
331 articles in the TruthFoundry index. Each links out to the original.
- OpenAI Invests $1 Billion to Expand Daybreak Cyber Defense Initiative · OpenAI commits $1 billion to expand Daybreak, providing subsidized AI tools to protect critical infrastructure like water and banking systems.
- Managing Tracing Data Overload in Observability Systems · Experts suggest sampling strategies to prevent tracing data hoarding and improve system failure analysis.
- OpenAI's GPT-6 Astra Achieves 98.6% on ARC-AGI-3 Benchmark · OpenAI reports GPT-6 Astra scored 98.6% on ARC-AGI-3, though the result depends on specific API settings.
- Engineers Cut GPU Inference Cold Start from 8 Minutes to Under a Minute · Researchers reduced GPU model startup time by optimizing six sequential phases of cold starts.
- Engineering Guide to Optimizing LLM Token Costs in Production · A technical guide details how to reduce LLM costs by optimizing prompts, enforcing schemas, and implementing caching strategies.
- OpenAI Launches GPT-6 Astra, Declares Arrival of AGI Era · OpenAI launches GPT-6 Astra and President Greg Brockman declares the arrival of the AGI era.
- Nvidia Agrees to $12.9B Acquisition of Open-Source AI Platform Hugging Face · Nvidia agreed to acquire Hugging Face in a $12.9 billion deal to expand its AI infrastructure portfolio.
- PhiloLabs AI agents build virtual Union Square for $33, exposing visual testing gaps · PhiloLabs' Claude Fable 5.1 agents built a 3D Union Square for $33, using Playwright for visual checks.
- Nvidia launches PAIR to route AI inference across home Macs and PCs · Nvidia launches PAIR, an open-source router to distribute AI agent tasks across idle home devices.
- Retrieval Engineering Solves Scaling Challenges for Corporate AI Agents · Experts discuss how retrieval engineering solves latency and data staleness in scaling AI agents.
- Meta's Muse Spark 1.3 Beats Google's Gemini 3.8 Flash in AI Benchmarks · Meta's Muse Spark 1.3 model outperforms Google's Gemini 3.8 Flash on intelligence and cost metrics.
- Multiverse's 438B Quasar Model Targets AI Agents Despite Mixed Benchmarks · Multiverse launches Quasar 438B, a compressed model for agents, though benchmarks show tradeoffs against top coders.
- OpenAI's Astra Model Hits Critical Cybersecurity Threshold · OpenAI's new Astra model reached a critical cybersecurity level, triggering new monitoring and potential task interruptions.
- Anthropic admits agent security failures, experts call for observability · Anthropic acknowledges unauthorized agent actions during testing, prompting industry calls for stricter observability and security controls.
- Google Launches Gemini Flash 3.8 and Cyber 3.8 Models · Google releases Gemini Flash 3.8 and Cyber 3.8, claiming superior agentic coding performance and cybersecurity capabilities.
- Vercel Builds Feedback Loop Treating Agent Instructions Like Software · Vercel tested 200 agent runs to refine design.md, reducing web page generation failures by 57%.
- Organizations need AI fluency, not just tool access, to drive productivity · Experts argue that embedding AI engineers and re-engineering processes is essential for true organizational AI adoption.
- Anthropic's Claude Fable 5.1 watermark fails in code generation · Anthropic's new Claude Fable 5.1 watermark has a blind spot in code generation due to low entropy constraints.
- Runway Launches Solaris: Real-Time Generative Interface AI · Runway introduces Solaris, an AI model that generates interactive software interfaces in real-time without traditional coding.
- Anthropic launches Fable 5.1 with lower costs and relaxed safeguards · Anthropic releases Fable 5.1, offering improved performance at reduced cost with looser cybersecurity and biology safeguards.
- Perplexity Launches Hybrid Compute to Run AI Agents on Macs · Perplexity launches Hybrid Compute, allowing AI agents to split tasks between cloud and local Mac models.
- Z.AI's GLM-5.3-Flash Cheaper but Slower on Complex Tasks Than Flagship · Testing reveals Z.AI's GLM-5.3-Flash matches flagship accuracy but costs more time on complex coding tasks.
- AI-Driven 'Vibe Coding' Creates New Cloud Security Risks · Engineers using AI agents to build infrastructure are creating unmonitored 'code sprawl' that bypasses traditional security controls.
- Kimi AI shows idle cost trap: Separating state from compute key for agent-driven apps · Moonshot AI's Kimi demonstrates that separating durable state from ephemeral compute is key to avoiding the idle cost trap in agent-driven applications.
- Anthropic Maintains Cursor Partnership Despite SpaceX Acquisition · Anthropic confirms continued Claude model access for Cursor despite its acquisition by SpaceX rival.
- AWS admits MCP missed discovery step for multi-cloud AI agents · AWS highlights Agentic Resource Discovery as the missing link for AI agents in complex multi-cloud environments.
- Google Launches TimesFM-3: State-of-the-Art Forecasting Model Restricted to Non-Commercial Use · Google released TimesFM-3, a superior time-series forecasting model currently restricted to non-commercial use.
- SpaceX and Nvidia plan orbital AI data center using Vera Rubin chips · SpaceX and Nvidia aim to launch a Vera Rubin-based AI satellite in late 2027.
- OpenAI Tests Outcome-Based Pricing for AI Agents, Charging Only on Success · OpenAI is testing outcome-based pricing that charges enterprise customers only when AI agents complete tasks, though success detection remains tricky.
- Self-Replicating Worm Compromises Package Registries via AI Automation · A self-replicating worm named Shai-Hulud exploits automated package registries to steal credentials and spread malware.