Loading Now

Representation Learning’s Grand Tour: From Universal World Models to Hyperbolic Genes and Trustworthy AI

Latest 40 papers on representation learning: Sep. 19, 2026

Representation learning continues to be a foundational pillar in AI/ML, enabling machines to understand, predict, and interact with the complex world around us. From deciphering the physical laws governing our universe to ensuring the trustworthiness and safety of AI systems, recent research showcases an incredible breadth of innovation. Let’s dive into some of the most exciting breakthroughs.

The Big Idea(s) & Core Innovations

A major theme emerging from recent work is the push towards unified, generalized, and domain-agnostic representation learning. A groundbreaking example is JEPA-Anything: Learning Predictive Models across Different Worlds by Taoyong Cui et al. from Shanghai Jiao Tong University and PHIAI Labs. This paper introduces Orthogonal Predictive Factorization (OPF), extending Joint-Embedding Predictive Architectures (JEPA) by decomposing latent targets into orthogonal factors. This allows a single predictive core to learn world models across astonishingly diverse domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and even weather. The key insight is that OPF provides a common learning principle to organize predictive states across heterogeneous domains, leading to scientific discoveries like validating biological interventions and recovering Keplerian scaling exponents from learned orbital modes.

Another significant development addresses challenges in multimodal and training-free learning. In Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning, Xinran Liu et al. from Nanjing University tackle semantic perspective misalignment in frozen multimodal large language models. They propose Semantic Perspective Anchoring, where task-required perspectives are explicitly associated with readout phrases. This allows extracting task-directed representations without parameter updates, bypassing the problem of models focusing on salient content rather than task-specific semantics. This training-free approach achieves significant gains on the MMEB benchmark and transfers across different model families.

In the realm of robotic manipulation, two papers highlight innovative approaches to creating robust 3D-aware and action-preserving representations. LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation by Wenbo Li et al. from South China University of Technology introduces a framework for learning 3D-aware scene representations from multi-view supervision and completing them from single RGB streams and observation history using generative completion with rectified flow. Their key insight is that memory and anchor mechanisms provide complementary benefits: memory handles occlusions, while anchoring prevents stale representations, leading to robust robotic control. Complementing this, ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models by Shijie Lian et al. from Huazhong University of Science and Technology, proposes a novel action tokenization approach that preserves physical relationships among robot actions during compression. They introduce Physical Rank Consistency (PRC), a metric that correlates more strongly with policy success than traditional reconstruction fidelity, achieved through Physical Rank Preservation and Quantization Regularization.

Addressing critical challenges in safety and trust, Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint by Zhiyun Jiang et al. from Sichuan University, introduces a novel task: visual scene negative captioning. Their CRCD framework reconstructs expected safe scene prototypes and contrasts them with reality to describe meaningful absent elements that pose safety risks. This tackles affirmation bias and limited mental filling capabilities in current vision-language models by mimicking human cognitive processing of negation. Furthermore, for enhancing security, Channel-Informed Neural Network for Physical Layer Key Generation by Jose Angel Sanchez Viloria et al. from Florida Atlantic University, learns reciprocity-preserving binary features directly from raw IQ measurements for physical-layer key generation, grounded in multipath channel structure. This allows high unique-key rates while maintaining security against eavesdroppers.

Medical and genomic applications also see significant strides. Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates by Yuhei Fujioka et al. (Cancerscan Inc., Kyoto University) proposes a unified pre-training framework for medical representation learning using Hierarchical Sub-token Aggregation (HSA), Partial Masking (PM), and Cross-Reference (CR). This effectively captures medical code hierarchies and diagnosis-treatment interactions, leading to data-driven rediscovery of promising drug repositioning candidates for Alzheimer’s disease. For genomics, HyCoSeq: Contextual Hyperbolic Representation Learning for Genomic Sequences by Chenhao Zeng et al. from ShanghaiTech University, shows that hyperbolic geometry provides a natural inductive bias for hierarchical genomic structures. By combining multi-curvature Lorentz convolution with bidirectional LSTM, HyCoSeq achieves competitive performance with models 2,500 times larger, demonstrating the power of geometry-consistent residual aggregation and sequence-level contextualization for compact models.

Finally, the theoretical foundations of representation learning are being re-examined. SL(n) Representation Learning: An Intrinsic Mixed-Curvature Space with Higher Curvature Capacities and Deeper Order-Aware Composition by Xingrun Li et al. from The University of Tokyo and Stanford University, introduces SL(n) as a novel representation space that intrinsically couples negative, zero, and positive flag curvature, enabling deep order-aware composition through nonzero nested Lie brackets at arbitrary depth. This work demonstrates that rich geometric structure can emerge from simple constraints, leading to improved performance on various graph benchmarks.

Under the Hood: Models, Datasets, & Benchmarks

These advancements are powered by innovative models, novel datasets, and rigorous benchmarks:

  • JEPA-Anything: Introduced a unified framework evaluated across diverse domains including vision, biology, clinical, control, molecular dynamics, physics, and weather. Publicly available Jepa-Anything Models and code.
  • Lens: Leverages existing frozen multimodal LLMs like Qwen2.5-VL and Qwen3.5, evaluated on MMEB (Massive Multimodal Embedding Benchmark), which includes 36 datasets for classification, VQA, retrieval, and grounding.
  • LIFD: Employs a memory-conditioned rectified flow model and GeoVGGT-R (a LoRA-adapted VGGT encoder). Evaluated on major robotics benchmarks like LIBERO, MetaWorld, and RoboTwin 2.0, and real-world UR5e tasks. The Scene Contradiction Index (SCI) is a novel diagnostic.
  • Pretrained Medical Representations: Utilizes large-scale Japanese medical claims and insurance data (5.5M+ individuals) with ICD-10 and ATC codes. Introduces Task-Adaptive Representation Approach (TARA).
  • Absence is Presence: Builds upon the SNUS dataset (12,118 instances) and uses LLaVA-1.5-7B, Qwen2.5-VL-3B, and Grounded SAM as base models. The CRCD framework is a key innovation.
  • ActionPiece: Introduces Physical Rank Consistency (PRC) as a new metric. Benchmarked on LIBERO, LIBERO-Plus, SimplerEnv, and VLA-Arena L0-L2. Related resources available at deepcybo-physai.github.io/ActionPiece.
  • Hyperbolic Graph Representation Learning: Constructs a patient-integrated biomedical knowledge graph using DOID, MONDO, HPO ontologies, and patient profiles from the Phenopacket Store. Evaluates HGCN and AttH models.
  • MyoFlow: A discriminative flow-matching framework evaluated on Hyser PR Dynamic and CEMHSEY datasets for HD-sEMG gesture recognition.
  • FreqSpaNet: Uses Spatio-Frequency Polarization Fingerprints (SFPFs) for hardware integrity detection. A dual-branch encoder with adaptive fusion and complementary pretraining is proposed.
  • LWVIC4Code: Builds on VICReg for non-contrastive code representation learning. Evaluated on Kamino (78,771 Python Type-IV clone pairs) and GPTCloneBench (Python, Java, C#, C Type-IV clones).
  • HyCoSeq: Multi-curvature Lorentz convolution with weighted Lorentzian residual aggregation and bidirectional LSTM. Benchmarked on TEB, GUE, and Genomic Benchmarks (GB) datasets. Code available at HyCoSeq.
  • Hub-Spectral Activation (HSA): A closed-form method for activating latent multimodal knowledge in frozen ImageBind and LanguageBind backbones. Code available at github.com/Luo1Yan/HSA.
  • FLAT: A unified multimodal encoder with text-to-image and image-to-text decoders, using nested dropout. Benchmarked on MS-COCO and Flickr30K for retrieval and generation. Project website: guangyusun.com/flat-website.
  • LM-PCVMNet: A deep learning framework for pediatric CVM staging, releasing PCVM+ dataset (1,800 lateral cephalometric radiographs with CVM stages, landmarks, and metadata). Code: github.com/ybupengwang/LM-PCVMNet.
  • DCRA: Repurposes forward diffusion as a corruption scheduler. Evaluated on CHB-MIT Scalp EEG Database for seizure detection. Encoder-agnostic, works with Mamba and Transformers.
  • CoMA-DiT: A bidirectional cross-modal Diffusion Transformer. Achieves SOTA on AVGC (auditory attention decoding) and DEAP (emotion recognition) tasks.
  • MUtE: Framework for concept erasure and counterfactual generation validated on synthetic data and NLP benchmarks like Bias in Bios, DIAL, Jigsaw. Code: github.com/toinesayan/MUtE.
  • DGCPath: Integrates diffusion-based view generation with distribution-aware variational contrastive learning. Evaluated on large-scale Aalborg, Chengdu, and Harbin traffic datasets. Code: github.com/Sean-Bin-Yang/DGCPath.
  • TripleBound: Hybrid framework for microservice decomposition, combining heterogeneous graph neural networks with parser-inferred triplet constraints. Evaluated on AcmeAir, DayTrader, PlantsByWebSphere, and JPetStore. Code: github.com/Mono2Distributed/triplet-guided-chgnn.
  • TailProp: A hierarchical vision backbone combining Gaussian and Cauchy propagation. Evaluated on ImageNet-1K, MS COCO 2017, ADE20K for classification, detection, and segmentation, as well as robustness and restoration tasks.
  • Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning: Focuses on Self-Knowledge Distillation, Semi-hard Outlier Sampling, and Outlier-aware Supervised Contrastive Learning. Evaluated on TinyImages-300K, SC-OOD, MOOD, and SYN benchmarks. Code: github.com/hyunjunchhoi/Three-factors.
  • CMA-OT: Utilizes a pre-trained multi-scale music expert (Jukebox) for hierarchical supervision. Achieves SOTA on AIST++ and TikTok datasets for dance-to-music generation. Project website: beria-moon.github.io/CMA-OT.
  • MCRL2: Multi-resource cross-attention-based representation learning for cloud microservice scheduling. Evaluated on Alibaba cluster traces v2021. Code: github.com/igeng/MSSchedRL.
  • PhysioAI: Clinical knowledge-guided semantic supervision for skeleton-based physiotherapy action recognition. Evaluated on KiMoRe, UI-PRMD, and Hard-67 stress test datasets. Uses CLIP text encoder and GPT-4o for CKD generation.
  • FRIST: A two-stage framework leveraging fMRI to guide EEG representation learning for individual finger decoding. Utilizes EEG dataset (https://doi.org/10.1184/R1/29104040).
  • GUIDE: LLM-driven preference elicitation framework. Evaluated using National Financial Capability Study (NFCS) Investor Survey and model portfolio data. Achieves 90%+ lower regret by turn 5. arxiv.org/pdf/2609.12137.
  • DGCPath: Self-supervised framework for path representation learning with diffusion-based view generator and variational contrastive learning. Code: github.com/Sean-Bin-Yang/DGCPath.
  • FLAT: A unified framework for cross-modal retrieval and generation with linearly interpolatable embeddings. guangyusun.com/flat-website.
  • Sanity Checking Causal Representation Learning: A real-world sanity check using a controlled optical experiment (light tunnel) with known ground-truth causal factors. Datasets and code: github.com/juangamella/causal-chamber, github.com/simonbing/CRLSanityCheck.

Impact & The Road Ahead

The implications of these advancements are profound. The ability to learn universal world models that generalize across diverse domains, as demonstrated by JEPA-Anything, hints at a future where AI systems can acquire understanding far more efficiently, perhaps even uncovering new scientific principles. The training-free approach of Lens means we can unlock the vast knowledge within existing frozen models without costly fine-tuning, accelerating deployment and accessibility of powerful multimodal AI. In robotics, improved 3D scene memory and physically-aware action tokenization pave the way for more robust and reliable autonomous agents capable of complex manipulation in unstructured environments. The integration of clinical knowledge into medical AI for drug repositioning and physiotherapy action recognition promises more personalized and effective healthcare solutions.

Critically, the field is also grappling with the trustworthiness and robustness of AI. The work on visual negative event understanding and physical layer key generation underscores the growing importance of safety and security in AI systems. The stark findings from the causal representation learning sanity check serve as a vital reminder: theoretical promises must be rigorously tested against real-world complexities. This will drive future research towards methods that are not just performant, but also robust, interpretable, and truly aligned with human understanding.

Looking forward, we can expect continued convergence of generative and discriminative models, the exploration of richer geometric spaces beyond Euclidean, and a stronger emphasis on physics-informed and knowledge-guided learning. The goal remains clear: to build AI that not only processes information but truly comprehends the world, enabling intelligent agents to learn, adapt, and operate safely and effectively across all domains of human endeavor. The journey of representation learning is more exciting than ever!

Share this content:

mailbox@3x Representation Learning's Grand Tour: From Universal World Models to Hyperbolic Genes and Trustworthy AI
Hi there 👋

Get a roundup of the latest AI paper digests in a quick, clean weekly email.

Spread the love

Discover more from SciPapermill

Subscribe to get the latest posts sent to your email.

Post Comment

Discover more from SciPapermill

Subscribe now to keep reading and get access to the full archive.

Continue reading