Towards Data Science
287 articles in the TruthFoundry index. Each links out to the original.
- Software Development Shifts from Using AI to Hiring Autonomous Agents · AI coding agents are evolving from tools to hired employees requiring new management structures and attention-based workflows.
- Three Moves to Maintain Judgment Amidst AI Hype · Experts advise product managers to understand the AI value chain and build foundational knowledge to resist vendor hype.
- AI Agents Generate CUDA Kernels, But Benchmarking Flaws Persist · AI agents generate CUDA kernels faster than PyTorch, yet flawed benchmarks and PyTorch's torch.compile complicate performance claims.
- Time Series Diffusion Fixes Gaussian Forecasting Uncertainty Flaws · This article explains how diffusion models resolve limitations in standard time series forecasting.
- New AI Roles Emerge as Automation Reshapes Data Science Careers · Emerging roles like AI Engineer and Forward Deployed Engineer redefine data science and machine learning careers.
- Core Flaws In Marketing Mix Models And Proposed Simulation Testing · This analysis documents limitations of modern Marketing Mix Models and validation methodology.
- Reinforcement Learning Explained: Origins, Key Components, and Challenges · Reinforcement learning enables agents to learn optimal behaviors through trial and error in dynamic environments.
- New Tool Measures Reliability of Non-Deterministic AI Coding Agents · Researchers introduce Coding Agent Consistency (cca) to measure LLM reliability in coding and analytics tasks.
- PINNs Outperform Finite Differences in 5D Quantum Problems · A 5D quantum experiment shows PINNs beat classical methods due to the curse of dimensionality affecting grids.
- AI-Assisted Learning Framework For Faster New Topic Mastery · Author outlines AI workflow methods to accelerate learning any new subject.
- Google Study Validates Spec-Driven Testing, Leaves Independent SDD Unmeasured · Google research confirms spec-driven test gains, author argues for independent agent separation.
- World Model Sync Cuts Drone Cloud Bandwidth Usage By 94% · Tutorial demonstrates shadow tracking method to reduce connected machine cloud data transfer.
- Four AI assistants tested on hidden traps in retail forecasting task · An experiment tested four leading AI models on a forecasting task with hidden production pitfalls.
- SIFT Computer Vision Algorithm Technical Workflow and Design Explained · This educational article breaks down the scale-invariant SIFT computer vision feature detection algorithm.
- Build Low Cost Reliable AI Model Routing Using Jev · This tutorial demonstrates building affordable reliable AI model routing with the Jev model.
- AI Agent Governance Must Evolve To Address Enterprise Agent Sprawl · Enterprises face growing AI agent sprawl requiring scaled centralized governance infrastructure.
- Reversal Curse fact recall failure replicated in tiny simple language models · Language models cannot reliably reverse learned factual associations, even at small scales.
- Research Measures LLM Agent Creativity On ML Engineering Tasks · Study evaluates LLM agent creativity using psychology frameworks on Kaggle ML tasks.
- Physics-Informed Neural Network Solves Navier-Stokes Blood Flow Problem · Custom PyTorch PINN reconstructs artery flow fields from sparse noisy velocity measurements.
- Separating Agent Development Lifecycle from Application Design · The author argues for coordinating agent development separately from the application it powers to manage tightly coupled loops.
- Building a Control Plane to Govern AI Agent Actions · A new control plane architecture is required to separate AI agent capabilities from authorized actions.
- Autoencoders Failed to Beat PCA in Anomaly Detection Tests · Experiments show default autoencoders do not outperform PCA on nonlinear anomaly detection tasks.
- Fewer LLM Calls Still Find Good Apartment Matches · An agent using a database query and cheaper models found 101 apartment matches with 25x lower cost than sending all listings to a strong model.
- ReLU Adoption Prioritized Training Speed Over Biological Plausibility · The shift to ReLU activation functions improved deep network training despite having less biological resemblance than sigmoid functions.
- Companies Cut AI Bills by Optimizing Token Usage and Caching · Businesses reduce AI costs by managing token consumption, leveraging caching, and selecting appropriate model tiers.
- Data Visualization Choices Shape Interpretation of the Same Dataset · Different chart types, scales, and aggregations of identical data lead to distinct viewer conclusions.
- Data Scientists Must Prioritize Inquiry Over Code Implementation · The author argues data science should focus on inquiry and results rather than code implementation, especially with coding agents.
- DAX Nested Measures Overwrite Filters: Solutions and Workarounds · DAX nested measures overwrite filter sets, requiring KEEPFILTERS() or separate measures to retain distinct values.
- Spec-Driven Test Automation: AI Agents Enforce Separation of Duties · A new workflow uses independent AI agents to write tests without seeing code, ensuring verification separation.
- Apache Iceberg Compaction Test Reduces 1000 Files To 6 For Performance Benchmarking · Author tests Apache Iceberg file compaction impact on query performance with local test data.