takara.ai
454 articles in the TruthFoundry index. Each links out to the original.
- Researchers Introduce ECHO Dataset for Contextual Object Placement · New synthetic dataset ECHO benchmarks AI agents on predicting object destinations using scene scans and human routines.
- PalmSpace Enables Precise On-Palm Interaction via Unified Touch Modeling · PalmSpace wrist-worn system enables precise, mode-aware input on the bare palm without user calibration.
- PureVision Framework Improves Medical VLM Lesion Interpretation · PureVision framework enhances multi-phenotype lesion interpretation in medical vision-language models using geometry supervision.
- AdSpark Dataset and Benchmark for Product-Centric Ad Video Generation · Researchers introduce AdSpark, a large-scale dataset and benchmark for generating product-centric advertisement videos.
- Gradient-Based Trajectory Optimisation for Sparse-View Cone-Beam CT · Researchers apply gradient ascent to optimize scanner poses for sparse-view CT scans.
- RT-DETR-World Transfers LLM Semantics to Real-Time Detection · Researchers propose RT-DETR-World, a compact detector transferring rich LLM semantics for efficient open-vocabulary detection.
- New Framework Achieves State-of-the-Art Zero-Shot Chinese Character Recognition · A global-to-local framework with radical verification achieves 83.06% accuracy on zero-shot Chinese character recognition.
- Researchers Introduce ScribbleEdit Benchmark for Image Editing · New ScribbleEdit benchmark reveals VLM/LLM models fail at scribble-only image editing.
- CRT-HMAR Framework Enables Open-Task Infrared-Visible Image Fusion · CRT-HMAR uses causal tracing and multi-agent regulation to generalize infrared-visible image fusion to unseen tasks.
- VIS-Ground System Enables Contextual Grounding in Interactive Video · VIS-Ground introduces a new framework for contextual grounding in interactive video generation systems.
- PhyDiCT Framework Reconstructs 3D CT Images from Sparse X-Rays · PhyDiCT uses differentiable rendering and diffusion priors to reconstruct 3D lung CT images from sparse X-ray projections without training.
- New 3D scene generation method works from single indoor or outdoor images · Researchers present improved single-image 3D scene generation for all environments
- ALIVE Framework Enables Interactive Object Insertion in Video Editing · ALIVE framework uses AI to make inserted video objects interact coherently with existing scenes.
- CtrlCache Accelerates Interactive Video World Models via Control-Aware Caching · CtrlCache framework speeds up interactive video generation by adapting computation to control sequences.
- PDB method introduced for improved facial animation retargeting · Researchers present PDB, a new artifact-free facial animation retargeting technique.
- New RGE framework improves image forgery detection using latent model knowledge · Researchers introduce Reserve-Guided Elicitation for accurate image forgery detection
- Anatomy-aware neural method improves CT-ultrasound medical image registration · Researchers present a deformable registration framework for CT and ultrasound medical imaging.
- LipDA Framework Detects and Attributes LipSync Forgeries · Researchers propose LipDA, a framework that detects and attributes LipSync forgeries by analyzing inconsistencies between lip motions and head poses.
- New AI Pipeline Improves Fringe Projection Profilometry Accuracy · Researchers introduce an AI-assisted pipeline to enhance Fringe Projection Profilometry characterization under challenging imaging conditions.
- ReGraph Model Reveals Dual Visual Streams Drive Relational Mapping · ReGraph model shows context-invariant codes emerge in dorsal visual streams before the hippocampus.
- S2PD diffusion method improves consistent video generation performance · Researchers introduce S2PD diffusion for physically consistent video generation
- Researchers Propose MTOR for Detecting AI-Generated Video · New MTOR framework combines multimodal semantics and temporal over-regularity to detect AI-generated videos.
- MoCAR Framework Achieves Top-Tier Trajectory Forecasting on Argoverse Benchmarks · MoCAR uses a continuous motion-code space to enable autoregressive trajectory forecasting without re-tokenization.
- CLIP benchmarked for zero-shot face and periocular gender estimation · Researchers evaluate CLIP performance for zero-shot gender estimation on face images.
- New Method Enables Casual Flash Lighting for Gaussian Splat Inverse Rendering · Researchers combine static and flash lighting to recover geometry and materials from casual indoor photographs.
- Research introduces UID normal form for nonlinear system state estimation · New research presents structural solution for nonlinear systems with unknown inputs
- Deepfake Detectors Fail as Generators Improve and Adversarial Attacks Emerge · Researchers found deepfake detectors losing accuracy and propose a new calibrated authentication paradigm.
- Dual-Rate Controller Reduces Force Error in Robotic Ultrasound · A model-based rate-limited controller reduces force error in robotic ultrasound imaging by accounting for image feedback delays.
- FlowHMR Framework Recovers Physically Plausible Human Motion from Video · FlowHMR uses flow matching and GRPO to generate physically trackable human motion from monocular video.
- VIDA-Geo: Multi-Agent System for Urban Intervention Design · VIDA-Geo is a multi-agent system that generates realistic urban interventions to improve city indicators like safety and greenery.