Loading Now

Foundation Models: From 4D Worlds to Humanoid Control and Medical Insights

Latest 100 papers on foundation models: Jul. 25, 2026

The world of AI/ML is buzzing with the transformative power of foundation models, pushing the boundaries of what’s possible across diverse domains. These large-scale pre-trained models, often adapting to new tasks with minimal data, are rapidly reshaping fields from robotics and healthcare to climate forecasting and molecular design. Let’s dive into some of the latest breakthroughs, exploring how researchers are harnessing and extending these powerful systems.

The Big Idea(s) & Core Innovations

One of the most ambitious endeavors is the creation of dynamic, physically plausible digital twins. Researchers at University of Massachusetts Amherst and Genesis AI introduce GS-Agent: Creating 4D Physical Worlds With Generative Simulation, a multi-agent framework that generates realistic 4D physical worlds from natural language by integrating physics engines. This agentic approach, decomposing complex tasks, significantly outperforms monolithic systems and ensures spatiotemporal consistency, a crucial step for embodied AI. Building on this, University of British Columbia and Johns Hopkins University present Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents, which automatically converts real-world robot interactions into simulatable digital twins. Their work highlights that open-source VLMs can achieve comparable results to proprietary models at significantly lower costs, democratizing scalable robotics simulation.

In the realm of robotics, several papers focus on robust, adaptive control. NVIDIA and Stanford University’s RoboTTT: Context Scaling for Robot Policies scales visuomotor context to an unprecedented 8K timesteps, enabling one-shot imitation and on-the-fly policy improvement by leveraging fast weights. This establishes context length as a new scaling axis for robot foundation models. Similarly, DAMO Academy, Alibaba Group, through RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model, introduces a family of embodied foundation models with native 3D grounding and contact-point prediction, demonstrating cross-embodiment generalization and the critical role of embodied pretraining for reasoning tasks.

Driving intelligence is also seeing significant advancements. Tsinghua University and AMap, Alibaba Group’s PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving tackles the prior-to-plan transfer problem in autonomous driving, preserving perception priors from frozen foundation models via adaptive expert routing. This leads to state-of-the-art performance with a single front-facing camera. For reliable spatial understanding in autonomous vehicles, Peking University and Wuhan University propose VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction, using visual and geometric cues from foundation models (DINOv2, VGGT) to improve 3D semantic occupancy prediction.

Healthcare applications are witnessing a paradigm shift. Hunan University’s MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning explicitly captures multi-scale temporal dynamics in EEG signals, achieving state-of-the-art performance across diverse brain decoding tasks. Building on multimodal approaches, University of Colorado Colorado Springs’s work on Multimodal Pretraining for Generalizable EEG Representation Learning jointly learns from raw EEG, time-frequency scalograms, and text embeddings for seizure detection, highlighting the persistent challenge of cross-subject generalization. For MRI analysis, University College London’s BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis utilizes a 3D Bi-Directional xLSTM-UNet architecture, demonstrating strong transfer performance across classification, segmentation, and regression. Similarly, University of Queensland’s SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens efficiently adapts SAM2 for 3D MRI segmentation, improving performance by integrating whole-volume context. In cardiac diagnostics, Maastricht University and University Hospital Münster leverage LLMs for automated label extraction and fine-tuned vision foundation models for disease classification from CMR images in their paper, Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images.

Beyond structured data, foundation models are being applied to complex, unstructured biological data. Stanford University introduces scVision: A vision foundation model for single-cell biology via spatial gene cartography, representing single-cell transcriptomics as images to achieve state-of-the-art zero-shot cell-type annotation and gene-program discovery. John Innes Centre demonstrates that Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes can detect antimicrobial resistance with high accuracy directly from metagenomic sequences using minimal probes on Evo 2 embeddings. For privacy-preserving single-cell analysis, Stanford University and Michigan State University propose TABULA: Predictive single cell foundation model for gene regulation and aging with privacy-preserving tabular learning, integrating federated learning and tabular learning for gene regulatory network recovery and rejuvenation factor identification.

Under the Hood: Models, Datasets, & Benchmarks

Recent research heavily relies on and contributes to a rich ecosystem of models, datasets, and benchmarks:

  • GS-Agent: Multi-agent framework with Manager, Entity, and Render agents. Leverages Genesis physics engine, BlenderKit, PolyHaven, and Meshy text-to-3D for asset generation. Outputs rich multimodal data beyond video.
  • WMFM-OOD: Framework for Out-of-Distribution detection in Wireless Multimodal Foundation Models for 6G ISAC. Utilizes geometric Base Station prototypes in a joint latent space. Evaluated on the DeepVerse6G dataset.
  • MSBraM: Self-supervised brain foundation model with multi-scale neural tokenizer and curriculum multi-scale masking. Pretrained on LaBraM pretraining data (18 public EEG datasets, 2,400+ hours). Evaluated on 12 public datasets including TUEV, TUAB, BCIC-2a, PhysioNet-MI, etc.
  • Multimodal EEG FM: Combines Mamba-based raw encoder, ViT-style time-frequency transformer, and retrieval-augmented text branch. Evaluated with Leave-One-Subject-Out (LOSO) on CHB-MIT scalp EEG dataset.
  • TransBiolab: New RGB-D dataset of 161,315 frames of cluttered transparent biomedical objects with 1.03M annotations, including 6D poses. Challenges existing foundation models like SAM 3 and FoundationPose. Released with 15 OBJ CAD models.
  • Aurora AI: Microsoft’s foundation model for atmospheric chemistry forecasting. Mechanistically interpreted using AuroraScope (sparse autoencoders) available at huggingface.co/hujason/aurorascope. Base model at microsoft/aurora/aurora-0.4-air-pollution on Hugging Face.
  • WSI-based Gene-Expression Prediction: Benchmarks deep regression models (Direct-ABMIL) and pathology foundation models (H-optimus-1). Evaluated on TCGA datasets (BRCA, HNSC, STAD, BLCA) and external cohorts like SCAN-B-Lund, SöS-BC-4, and KS-Solna.
  • PerceptDrive: Utilizes frozen foundation models as perception providers (geometric, semantic, dynamic priors). Evaluated on NAVSIM v1 and NAVSIM v2 benchmarks.
  • WMFM-OOD: Leverages geometric BS prototypes and temperature-scaled scoring. Evaluated on the DeepVerse6G dataset.
  • Time Series Foundation Models: Papers extensively use and compare models like TimesFM (Google), Chronos (Amazon), and MOIRAI (Salesforce). Applied to HRV forecasting from Samsung Galaxy Watch Active 2 data and crowd forecasting for SAIL2025 event data.
  • Air Quality Arena (AQA): A comprehensive dataset and benchmark for air quality forecasting. Includes 14,000+ station-pollutant series across 7 countries and 6 pollutants. Benchmarks 11 TSFMs, including cross-modal VisionTS++. Dataset and framework at AirQualityArena.github.io.
  • DBMol: Framework for de novo small molecule design, guided by differentiable signals from structure prediction models like Boltz-2 and AlphaFold-3. Uses a discrete flow-matching model (DeFoG) and SynFormer for synthesis-aware post-processing.
  • MASHT: Time series classification pipeline combining MultiRocket and Hydra random convolutional transforms with TabPFN-3 tabular foundation model. Code at github.com/joschac/masht.
  • IGGT4D: Streaming instance-grounded geometry Transformer. Introduces InsScene4D-147K dataset with 147K video sequences and geometry-consistent instance annotations.
  • Cognitive Dual-Process Planning: Utilizes perception foundation models like Grounding DINO, Segment Anything Model, and Depth Anything Model for structured scene knowledge. Employs Qwen3-VL-235B-A22B as an expert VLM.
  • Agentic Real2Sim: Converts DROID robot manipulation episodes into MuJoCo-simulated digital twins. Generalizes across rigid, deformable, and humanoid domains. Supports interchangeable VLM backends.
  • Latent Riemannian Flow Matching: Operates in the latent space of frozen geometric foundation models like VGGT. Reduces Chamfer distance on ScanNet++ and ETH3D. RGB decoder for novel view synthesis.
  • Delineate Anything v2: Global foundation model for field boundary mapping. Built on FBIS-73M, a 73-million-instance multi-resolution dataset spanning 61 countries. Code at github.com/Lavreniuk/Delineate-Anything.
  • THOR and TerraMind: Geospatial Foundation Models (GFMs) developed by ESA’s φ-lab. THOR unifies Sentinel-1, -2, -3. TerraMind is a multimodal generative model. Evaluated across ten Earth observation use cases using TerraTorch.
  • GigaPath-Flash and GigaTIME-Flash: Efficient pathology foundation models distilled from billion-parameter GigaPath. GigaPath-Flash uses ViT-S + LongNet. GigaTIME-Flash predicts spatial proteomics. Apache-2.0 licensed open weights available.
  • SAMRI-3D: Benchmark and method for 3D MRI segmentation, adapting SAM2 with Global Volume Tokens. Evaluated on SAMRI-3D benchmark (10,392 volumes from 34 datasets).
  • ChoroplethMap-Bench: First controlled benchmark for evaluating FM spatial understanding using 2,400 synthetic maps and 12,000 questions. Compares 22 frontier models.
  • RynnBrain 1.1: Family of embodied foundation models (2B, 9B, 122B-A10B). Benchmarked on VSI-Bench, MMSI, and RefSpatial-Bench. Code and models at alibaba-damo-academy.github.io/RynnBrain and huggingface.co/collections/Alibaba-DAMO-Academy/rynnbrain-11.
  • MEVION: Low-cost open-source dual-arm robot data collection system. Hardware designs, software, and STEP files are open-source at github.com/haraduka/mevion.
  • MuViSeg: Multi-view segment correspondence matching. Uses MASt3R and VGGT features. Evaluated on Replica and Virtual KITTI 2, and HM3D Instance Image Navigation.
  • AE-PSL: Communication-efficient parallel split learning with AutoEncoder compression. Fine-tunes ViT-B/32 on CIFAR-100, Food101, SUN397, FEMNIST. Code at github.com/Nousphera/AE_PSL.
  • Zero-Shot Crowd Forecasting: Evaluates TimesFM and Chronos-2 on SAIL2025 event pedestrian flow data.
  • LFM: Framework for Source-Free Universal Domain Adaptation leveraging CLIP (ViT-B/16) and GPT-4-turbo. Evaluated on Office-31, Office-Home, VisDA, and DomainNet. Code at github.com/iamjingli/LFM.
  • Pixel-Space Diffusion Transformers (pDiT): Survey categorizing architectures for latent-free image generation directly in pixel space.
  • A Weisfeiler-Leman Characterization of Global-Attention Graph Transformers: Theoretical and empirical analysis of graph transformers for MILPs. Uses Ecole library for MILP representations. Code at github.com/Abrar2652/optfm_probing/.
  • Reinforcement Learning: From Algorithms To Foundation Models: Dissertation presenting a unified view of RL, including Diffusion World Model (DWM) and Consistency Models as RL policies (CM-RL).
  • Lightweight Wrappers for Regional Drought Forecasting: SMR2 and MBB inference-time wrappers for TimesFM, Chronos, and MOIRAI. Evaluated on SPEI-30 forecasting. Code at github.com/Wentao-Gao/Lightweight_Wrappers.
  • RGMR: Residual-Guided Multi-Resolution Refinement for drought forecasting. Architecture-agnostic across TimesFM, TimeGPT, and TabPFN. Code at github.com/Wentao-Gao/RGMR_implementation.
  • Node4All: Node representation learner using Channel Graph Transformer (CGT) trained on synthetic graphs. Code at github.com/dooho00/node4all.
  • Lite-PolypInductor (Lite-π): Enhances lightweight polyp segmentation models (U-Net, PraNet) using SAM, DINOv2, OneFormer priors via UFMI module. Code at github.com/lostinrepo/Lite-Pi.
  • VIDAR: Hybrid visual-inertial dense reconstruction. Combines SVO+IMU with Depth Anything 3 (DA3). Evaluated on EuRoC and TUM RGB-D datasets.
  • DepthART: Transfers foundation monocular depth capabilities to tiny models. Uses TinyViM+DPT backbone with Bias-Resistant Data Sampling (BRDS) and Camera-Conditioned Fine-Tuning (CamFT).
  • HarmoHOI: Unified diffusion framework for multi-view hand-object interaction (HOI) synthesis. Integrates Mixture of Multi-view DiT (M2DiT) and Global Motion Aligning Diffusion (GloMAD). Project page: droliven.github.io/HarmoHOI_project.
  • CNS-Edit++: Latent-space 3D shape editing using Coupled Neural Shape (CNS) representation. Supports category-agnostic editing with TRELLIS and Direct3D-S2 foundation models.
  • RegionFM: Interpretable brain MRI classification using region-wise foundation model embeddings (BrainIAC, NeuroFM) and anatomical segmentation (FastSurfer). Evaluated on OASIS-1 dataset.
  • xperception: Zero-shot 6D pose estimation using CAD models and foundation models (DINOv2, GeDi). Based on the FreeZe algorithm.
  • It Depends on the Dataset: Investigates TRIBE v2 (Meta FAIR) and V-JEPA2 for video memorability prediction on Memento10k and VideoMem. Code for TRIBE v2 at github.com/facebookresearch/tribev2.
  • NeoST: First spatio-temporal foundation model pretrained solely on procedurally generated synthetic data. Evaluated on PEMS and METR-LA datasets.
  • OpenMHC: Largest open-access wearable health dataset (67M hours, 11,894 participants). Includes open-source reimplementations of Apple’s WBM and Google’s LSM-2. Dataset at myheartcounts.stanford.edu/openmhc.
  • M-CPT: Multimodal Continual Pre-Training for DINOv2, SigLIP, AIMv2. Uses Continual Position Embedding (CPE) and Alignment Loss with Qwen2.5 LLM. Evaluated on ChartQA, DocVQA, AI2D, MMMU.
  • PAMT: Prompt-guided Adaptive Model Transformation for pathology image classification. Uses Representative Patch Sampling and Prototypical Visual Prompts. Code at github.com/hust-linyi/PAMT.
  • Revisiting data-driven dynamic security assessment with a tabular foundation model: First application of TabPFN for pre-fault Dynamic Security Assessment (DSA) in power systems.
  • DPNeXt: Lightweight multi-scale feature fusion decoder for multi-task dense prediction. Uses frozen DINOv2-Reg backbone. Code at github.com/kangjehun/DPNeXt.
  • IMBench: Benchmark of 35 robotic manipulation tasks for ‘intuitive manipulation’. Compares Diffusion Policy, π0.5, and GR00T 1.5. Uses Robosuite.
  • EpiDistill: Geometric distillation framework for monocular depth, transferring priors from multi-view foundation models. Improves UniDepthV2, DepthPro on ETH3D, DIODE, KITTI, NYU, ScanNet, etc.
  • Are All Tokens Necessary for Visual Place Recognition?: Benchmarks token reduction methods for VPR using Vision Transformers (DINOv2). Code at github.com/Tong-Jin01/TokenReduction4VPR.
  • Recursive Harness Self-Improvement (RHI): Prompt-level harness optimization for agents. Evaluated on sonnet-4.6, opus-4.7, opus-4.8 models.
  • LLM4EHR: Multimodal fusion framework aligning clinical time series with medical event sequences via domain-adapted LLMs. Evaluated on MIMIC-IV and Physionet Challenge 2012 datasets. Code at github.com/CrankyWilliam/LLM4EHR.
  • HDR (Hierarchical Denoising for Visual Reasoning): Unified framework for multi-step visual reasoning using streaming autoregressive diffusion. Project page: hierarchical-diffusion-reasoning.github.io.
  • Motion-Conditioned Multi-View Fusion: MCF-Net combines EchoPrime foundation model with sparse motion priors for myocardial infarction localization. Uses HMC-QU dataset and CoTracker3.
  • Symbal: Detects systematic misalignments in MLLM-generated captions. Uses foundation models to identify textual errors and associated visual features. Introduces SYMBALBENCH with 1.7M image-text pairs. Code at github.com/Stanford-AIMI/Symbal.
  • Scaling Behavior Foundation Model for Humanoid Robots: Uses Humanoid Transformer. Leverages LAFAN, AMASS, OMOMO, GRAB, SnapMoGen, FineDance, BONES-SEED, Embody3D datasets.
  • SUFLECA: Weakly-supervised zero-shot CAD-to-image alignment method. Scales NOC-supervised feature learning on 674K images from 12 datasets. Code at github.com/snt-arg/SUFLECA.
  • ViPS: Harmonizes diverse visual priors from multiple foundation models (VGGT, DepthAnything3, TraceAnything, Wan2.1, RADIO) for MLLM spatial understanding. Code at visual-ai.github.io/vips.
  • CFM-Bench: Unified multi-domain, multi-task benchmark for Channel Foundation Models (CFMs). Integrates DeepMIMO O1 60, Sionna, MOCSID, DICHASUS ADXX, MaMIMO-UAV, and Multimodal-Wireless datasets.
  • DynaBase: Minimal two-parameter model for zero-shot dynamical systems reconstruction, iteratively reduced from DynaMix and Chronos.
  • WanSong v1.0: Pure diffusion-based music foundation model generating high-fidelity, multilingual songs. Uses a custom VAE and Hybrid-MMDit transformer. Project report: arxiv.org/pdf/2607.14749.
  • Pretraining Multiple Instance Learning Networks: Distillation-based pretraining using pathology slide foundation models (TITAN, CARE). Evaluated on TCGA-UT-8K, BCNB, BRACS, CPTAC, EBRAINS, KIDRARE, MUT-HET-RCC. Code at github.com/fu0201/MIL_Pretrained.
  • CERPE: Communication-efficient relative pose estimation for multi-robot systems. Coordinates VGGT, Metric3Dv2, SALAD. Evaluated on CARLA simulator and OpenMars dataset. Project page: rokeeeto-li.github.io/cerpe.github.io/.
  • GCA-Bench: Benchmark for complex robotic grasping (102 tasks). Evaluates AnyGrasp, GraspMAS, OpenVLA, π0, π0.5. Project page: airvlab.github.io/GCA-Bench/.
  • XCT-SAM: Two-stage parameter-efficient adaptation of SAM for industrial XCT defect segmentation using Conv-LoRA. Evaluated on synthetic OoD and real NIST XCT benchmarks. Code at github.com/Mahedi-61/XCT-SAM.git.
  • SeeSE3: Investigates 3D space emergence in vision features (DINOv2) using ‘Poincare Adapter’. Evaluated on ScanNet, ARKitScenes, TUM RGB-D, 12 Scenes, 7 Scenes.
  • TEDDY: Pediatric Foundation Model for Risk Forewarning (1.84M params) trained on 73M ICD-10 diagnoses from 1.6M children.
  • RxBrain: Embodied Cognition Foundation Model with joint language-visual reasoning and imagination. Uses Unified multimodal Mixture-of-Transformers (MoT) architecture. Hugging Face model: huggingface.co/tencent/Hy-Embodied-RxBrain-1.0. Code at github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0.
  • UniMedSeg: Transformer-centric foundation model for 2D/3D medical image segmentation. Uses Decoupled Split Attention. Code at github.com/Lii1228/UniMedSeg.
  • MBTI: Multi-Branch Efficient Fine-Tuning for Hyperspectral Image Classification. Uses HyperSIGMA foundation model. Evaluated on KSC, Salinas, FangluTeaFarm datasets.
  • RFMSR: Residual Flow Matching for Image Super-Resolution. Vision-only framework. Code at github.com/Faze-Hsw/RFMSR.
  • FM2: Unified federated learning for heterogeneous multimodal medical imaging. Uses dual Mixture-of-Experts and GPT-4o for caption generation. Evaluated on MIMH benchmark, SLAKE, VQA-RAD, VQA-Med.
  • OptCar: Specializes generalist forward kinodynamic (FKD) prediction foundation model (AnyCar) for high-speed off-road vehicle control. Project page: amrl.cs.utexas.edu/optcar.
  • Tabular Foundation Models for Discrete Choice Estimation: Applies TabPFN to discrete choice problems.
  • Cost-Optimal Foundation Model Deployment Portfolio: Uses GPT-4o, GPT-4o-mini, Gemini-2.5-Flash, Claude Haiku 4.5, Llama-3.1-8B, Qwen2.5-VL-7B, InternVL2-8B, YOLOv8-L.
  • SPINE: Agentic AI framework for robot debugging. Evaluated on DOBOT X-Trainer and AgileX PiPER platforms.
  • The Spectrum Is Not Enough: Analyzes TimesFM and Chronos-Bolt for time-series forecasting. Code at github.com/KurbanIntelligenceLab/SINE.
  • MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model. Jointly trained on Static State Estimation and AC Power Flow. Uses gridfm-datakit and PyTorch Geometric.

Impact & The Road Ahead

These advancements are collectively ushering in an era of more capable, adaptable, and efficient AI systems. The ability to generate realistic 4D physical worlds and convert real robot interactions into digital twins (GS-Agent, Agentic Real2Sim) is foundational for training embodied AI agents, reducing the simulation-to-reality gap, and accelerating robotics research. The emphasis on long-context processing (RoboTTT) and native 3D grounding (RynnBrain 1.1) promises robots that can understand and interact with complex environments more intelligently and autonomously.

In medicine, foundation models are poised to revolutionize diagnosis and personalized treatment. From hierarchical EEG analysis (MSBraM) to automated cardiac diagnosis (Maastricht University & University Hospital Münster) and the efficient analysis of whole-slide images in pathology (GigaPath-Flash and GigaTIME-Flash), these models promise to enhance diagnostic accuracy, reduce manual workload, and extend access to molecular profiling. The development of privacy-preserving techniques (TABULA, AE-PSL, FM2) is particularly crucial for unlocking the full potential of sensitive medical data through collaborative learning.

However, challenges remain. The insights from WMFM-OOD highlight the critical need for robust out-of-distribution detection in safety-critical systems like 6G. The struggle with cross-subject generalization in EEG (University of Colorado Colorado Springs) and the finding that vision encoders don’t inherently exhibit human-like color thresholds (Computer Vision Center, Spain) remind us that human-centric alignment is not a given. The imperative to move beyond purely predictive models to Mechanistic World Models (University of Oxford & MPI for Intelligent Systems) points towards a future where AI not only forecasts but also provides interpretable explanations, fostering genuine scientific discovery. The emergence of specialized yet efficient inference-time adaptation strategies (Lightweight Wrappers, RGMR, DepthART) further enhances the practical deployability of these powerful models across diverse, resource-constrained environments. The journey of foundation models is only just beginning, promising an exciting future of intelligent automation and discovery.

Share this content:

mailbox@3x Foundation Models: From 4D Worlds to Humanoid Control and Medical Insights
Hi there 👋

Get a roundup of the latest AI paper digests in a quick, clean weekly email.

Spread the love

Discover more from SciPapermill

Subscribe to get the latest posts sent to your email.

Post Comment

Discover more from SciPapermill

Subscribe now to keep reading and get access to the full archive.

Continue reading