Representation Learning Unpacked: From Physical Biases to Self-Supervised Horizons
Latest 75 papers on representation learning: Aug. 15, 2026
Representation learning, the art of transforming raw data into meaningful and useful numerical forms, remains a cornerstone of AI/ML innovation. It underpins everything from understanding complex biological systems to navigating autonomous agents and enhancing user experiences. The sheer diversity of data – from images and videos to patient records, financial transactions, and even molecular structures – presents a continuous challenge: how do we extract the right information in the right way? Recent research has pushed the boundaries, exploring novel approaches that leverage physical priors, self-supervision, and clever architectural designs to create more robust, efficient, and interpretable representations.
The Big Idea(s) & Core Innovations
At the heart of these advancements is a drive to imbue representations with richer, more relevant inductive biases, often inspired by the underlying domain or learned through sophisticated self-supervision. For instance, in visual tasks, a study titled “A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources” by Brunó B. Englert and Gijs Dubbelman from Eindhoven University of Technology highlights that DINOv2-style pretraining consistently delivers strong performance under limited resources. However, it also uncovers a fascinating trade-off: combining DINOv2 with video SSL objectives improves semantic understanding but can degrade geometric tasks. This suggests a need for targeted representation learning based on downstream task requirements.
Taking this further, in “Dual-Manifold Geometry Guided Representation Learning: Adaptive Coupling between Kernel and Data Spaces” by Wencong Zhang et al. from Southern Medical University, a dual-manifold perspective is introduced. Their Kernel-Guided Feature Transform (KGFT) module transfers geometric information from convolutional filters (kernel manifold) to reshape feature representations (data manifold), leading to consistent improvements in both CNNs and Transformers. This demonstrates that network parameters themselves hold valuable geometric priors that can be actively leveraged.
Uncertainty modeling is another crucial area. “Capturing Uncertainty in Human Motion for Representation Learning in Soccer” by Yizhou Xu et al. from KTH Royal Institute of Technology and EA Sports TRACAB introduces discrete distribution learning (DDL) to model probabilistic distributions over future motions, a critical step for understanding inherently multimodal human movement in complex environments like soccer. This moves beyond deterministic predictions to embrace the true nature of dynamic systems.
From a foundational perspective, Mathieu Cyrille Simon et al. from UCLouvain, EPFL, in “Unsupervised Disentanglement Without Compromises: How Functional Orthogonality Enforces Identifiability” challenge long-held beliefs about unsupervised disentanglement. They propose that defining latent concepts through functional orthogonality of a generative mapping’s Jacobian, rather than statistical independence, is the key to identifiability. This theoretical breakthrough could redefine how we approach building interpretable models.
Multimodal learning is seeing rapid evolution. The “Generation-Augmented Supervision for Multimodal Understanding (GAS)” framework shows that using visual generation as training-time auxiliary supervision (rather than an inference-time capability) significantly boosts multimodal understanding. Meanwhile, “Multiview Representation Learning via Distributed Joint Latent Space Structuring” by Milad Sefidgaran et al. from Huawei Paris Fourier Research Center proves that cross-view statistical correlations in representations actually tighten generalization bounds, a counter-intuitive finding that justifies feature alignment in distributed multiview settings. For medical data, “Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training” introduces AID, a framework modeling the intrinsic 2D hierarchical structure of tabular medical data with importance-aware masking and soft-label discretization for robust cross-domain generalization. This shows a deep understanding of data structure can yield significant benefits.
Other notable innovations include: * Relational Deep Learning: “Incremental Evaluation and Training in Relational Deep Learning” by Jakub Peleška and Gustav Šír from Czech Technical University in Prague demonstrates that incremental fine-tuning outperforms expensive retraining when dealing with prevalent temporal concept drift in real-world relational databases. * Fairness: Yijin Ni and Xiaoming Huo from Georgia Institute of Technology, in “A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes”, propose using a joint discrepancy measure (HSIC) for fair representation learning, achieving comparable fairness-accuracy tradeoffs while training significantly faster. * Physics-Informed Neural Networks: Yulun Wu et al. from KTH Royal Institute of Technology introduce FALM-PINN in “Alternating Levenberg-Marquardt Training of Physics-Informed Neural Networks with Fourier-Enhanced Features”, which decouples representation learning from coefficient fitting using Fourier-enhanced features, achieving two orders of magnitude lower errors on challenging PDEs by addressing spectral bias. * Federated Learning: “Sheaf-Based Federated Representation Learning” by Gabriele D’Acunto et al. from Sapienza University of Rome proposes SFRL, using learnable network sheaves to align heterogeneous agent-specific latent spaces without a shared global space, crucial for semantic communication. “Attention, Anomalies! Handling Attention Layers in Unsupervised Federated Outlier Detection” by Mihailo Ilić et al. from University of Novi Sad addresses specific aggregation techniques for Memory-Augmented Autoencoders (MemAE) in Federated Learning, demonstrating that guided clustering methods improve robustness in unsupervised anomaly detection, especially in non-IID settings.
Under the Hood: Models, Datasets, & Benchmarks
These advancements are often enabled by, or lead to, new models, datasets, and evaluation benchmarks. Here are some key examples:
- Causal World Models: “A Unifying Perspective on Causal World Models: From Observations to Representations to Structure” by Avinash Kori and Fabrizio Russo (Imperial College London) proposes a formal framework for CWMs as structured Markov decision processes, clarifying identifiability conditions for representation, transition, and causal structure components.
- Self-Supervised Visual Pretraining: “A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources” extensively uses the K700 (Kinetics-700) dataset for pretraining, and evaluates on ImageNet-1K, Pascal VOC, Cityscapes, ADE20K, NYUv2, KITTI, SSV2, RE10K, MOVi-F. Code is available at github.com/tue-mps/vision-ssl-study.
- RGB-Event Video Person Re-Identification: The “Paths: Prompt-aware Spatio-temporal Transformer…” framework by Yakun Huo et al. (Dalian University of Technology) achieves SOTA on EvReID, MARS, and iLIDS-VID datasets. Code: https://github.com/Reflection0427/Paths.
- Relational Deep Learning: “Incremental Evaluation and Training in Relational Deep Learning” leverages the RelBench benchmark (https://relbench.stanford.edu/) and existing PyTorch-based GNN/Frame libraries.
- Dual-Manifold Representation Learning: “Dual-Manifold Geometry Guided Representation Learning…” demonstrates improvements on CIFAR-100, ImageNet-1K, MATH10K, GSM8K, MAWPS, SVAMP, and AQuA datasets. Code: https://github.com/ZWC-SMU/KGFT.
- Analog Circuit Design: “Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization” utilizes Open Circuit Benchmark (OCB), Ckt-Bench-101, and Ckt-Bench-301 datasets.
- Multiview Representation Learning: “Multiview Representation Learning via Distributed Joint Latent Space Structuring” tests on CIFAR10, CIFAR100, and IMDB-WIKI datasets.
- Multimodal Understanding: The GAS framework leverages COYO-700M for T2I prompts and evaluates on CV-Bench, MathVista, BLINK, MME, MMMU, MVBench, ZebraLogic, MMLU-Redux, RefCOCO, ImageNet.
- Binary Code Representation Learning: “Instruction Alignment for Binary Code Representation Learning” by Huaijin Wang and Shuai Wang (Shandong University, HKUST) uses BinaryCorp and BinKit datasets to train and evaluate on instruction-level correspondences.
- Speech Processing: “Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization” uses W2v-BERT 2.0, Qwen2.5-0.5B and evaluates on SeedTTS test sets, LibriSpeech-PC-test-clean.
- 3D Scene Understanding: “STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding” by Mingwei Xing et al. (KE Holdings Inc.) is evaluated on ScanNet, S3DIS, Structured3D, nuScenes, Waymo, Matterport3D, ARKitScenes, 3D-Front, HM3D, SpatialLM. Code: https://github.com/KE-CS/STAR.
- Noisy Label Aggregation: “Dual-Primal Graph VAEs for Noisy Label Aggregation” by Patrick Stinson and Nikolaus Kriegeskorte (Columbia University) uses MNIST, CIFAR10, CIFAR10-N and various crowdsourcing benchmarks. Code: https://arxiv.org/pdf/2608.11473.
- Sign Language Processing: “Gloss-Free Representation Learning for Cross-Dataset Sign Spotting” introduces TSL-News (Turkish broadcast corpus) and TSL-SB (Turkish Sign Language Spotting Benchmark). They also utilize TSLD.
- Long-Tailed Classification: “CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification” by Gawon Lim (University of Illinois Urbana-Champaign) evaluates on CIFAR-100-LT, ImageNet-LT, and Places-LT.
- Reliability Estimation in Medical AI: “Decodable but Not Accessible: Auditing Distance-Based Reliability Estimation…” uses ISIC 2018 and PAD-UFES-20 datasets.
- Multimodal Intrinsic Dimension Estimation: “FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data” by Viktoria Schuster et al. (MIT) provides code at https://github.com/viktoriaschuster/FiGuRO.
- Audio Representation Learning: “DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning” uses FSD50K for pretraining and evaluates on ESC-50, Speech Commands v2, UrbanSound8K, GTZAN. Evaluation protocol: https://github.com/nttcslab/eval-audio-repr.
- Audio Effects Representation Learning: “Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning” by Xinlu Liu et al. (Tencent Music Entertainment) uses MUSDB18 and MoisesDB. Code: https://relative-fx.github.io.
- Text-Based Image Retrieval (TBIR): “Rethinking Text-Based Image Retrieval in Specific Domain” introduces SecMM-TBIR benchmark (50k surveillance images) and DSMM-TBIR data engine.
- Multimodal Medical Pre-training: “Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training” uses SLICE-3D, HOP, and EyePACS datasets. Code: https://github.com/Ethan-ysliu/AID.
- Fair Representation Learning: “A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes” provides code at https://github.com/Yijin911/FRHSIC.
- MRI Motion Artifact Reduction: “Motion Artifact-Aware Self-Supervised Representation Learning…” by Mojtaba Safari et al. (The University of Chicago) uses IXI, HCP, and MR-ART datasets.
- Federated Representation Learning: “Sheaf-Based Federated Representation Learning” provides code at https://github.com/SPAICOM/sheaf-based-federated-representation-learning.git.
- Explicit Primitives for Medical Imaging: “Implicit representations are dead. Long live explicit primitives!” by Nil Stolt-Ansó et al. (Technical University Munich) benchmarks on BACH histology and NLST lung CT datasets.
- Compressed Sensing: “Flow-Based Generative Modeling for Optimizing Sampling Policies…” uses MNIST, CelebA, and fastMRI knee datasets.
- 2D-3D Matching: “TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching” by Chongjian Wang and Junjie Gao (Shandong Women’s University, Shandong University of Science and Technology) uses 7-Scenes and RGB-D Scenes V2 datasets.
- Cardiac Anatomy Generation: “Flow-based conditional cardiac anatomy generation for virtual cohorts” utilizes UK Biobank data.
- Deepfake Detection: “Foundation Models are Implicit Deepfake Detectors” by Stefan Smeu et al. (Bitdefender) uses GenImage, FakeAVCeleb, AVLips, DeepfakeEval-2024, MAVOS-DD, MMDF, DFDC datasets, and models like DINOv3, AV-HuBERT, RAVEn, BRAVEn, PE-Core, BEiT, OpenCLIP, SigLIP 2.
- Invariant Learning: “From Objectives to What Models Learn: A Landau Theory of Invariant Learning” provides a theoretical framework for various invariant learning methods.
- Multimodal Federated Learning: “Multimodal Federated Learning under Dual-Axis Modality Missingness” uses PAMAP2, RealWorldHAR, SleepEDF, ADNI datasets. Code: https://github.com/AdibaOrz/Flux.
- Graph Sequence Models: “HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models” uses ECHO-SYNTH and LRIM-16 benchmarks. Code: https://anonymous.4open.science/r/HOPPER-0593/.
- Whole-Slide Pathology: “Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception” by Jingyun Chen et al. (Harbin Institute of Technology, The University of Hong Kong) uses SlideBench-VQA, WSI-VQA, PathVQA, TCGA cohorts, PathMMU. Code: https://github.com/G14nTDo4/AdaptivePath.
- Trajectory-User Linking: “Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking” by Zhifeng Chua et al. (Shanghai Normal University) uses Foursquare-NYC, Foursquare-TKY, Foursquare-JKT. Code: https://github.com/superior-DL/MakeTUL.
- RF Sensing: “FSTC-Encoder: Feature–Spatial–Temporal Correlation Learning for Generalizable RF Sensing” uses Widar3.0, CSI-Bench, XRF55 datasets.
- Influence Maximization: “Rethinking Learning-Based Influence Maximization: Simple Neural Surrogates and Native Discrete Search” by Yiqiao Liao and Parinaz Naghizadeh (UC San Diego) provides code at https://github.com/yl489/rethink-IM.
- Multivariate Time Series Classification: “FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework…” by Pingping Liu et al. (Jilin University) uses UEA Multivariate Time Series Classification Archive. Code: https://github.com/Wangmy2120/FreSH00.
- MALDI-TOF Mass Spectrometry: “Biologically Informed Representation Learning for Robust Cross-Center Generalization…” uses DRIAMS-A/B/C/D, MARISMa, RKI, MS-UMG datasets. Code: https://github.com/alexgaarciia/MALDIAlign.
- Multi-Omics Integration: “DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment…” uses TCGA pan-cancer cohort, METABRIC, MetaCancer, and an institutional colorectal-cancer survival cohort.
- Wearable Disentangled Representations: “Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations” by Ioannis N. Ziogas et al. (Khalifa University of Science and Technology, University of Toronto) uses the HARWE dataset. Code: https://github.com/GiannisZgs/OmniDecVAEs.
- Knife Image Retrieval: “KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement” introduces the KnifeHunter Dataset. Dataset: https://doi.org/10.5281/zenodo.19210078.
- Small-Data Time Series: “Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models…” uses a Cryocooler telemetry dataset.
- Corruption Robustness: “Suppress and Diversify: Refining Robust Pathways for Corruption Robustness” provides code at https://github.com/JGyoung-UCAS/suppress_and_diversify.
- Reinforcement Learning: “Flowing Through States: Neural ODE Regularization for Reinforcement Learning” uses Atari Learning Environment (ALE) and Minigrid environments.
- Electronic Health Records: “MiGHT-EHR: A Multi-task Graph Transformer for Heterogeneous Temporal Electronic Health Records” by Anirudh Rayas et al. (Arizona State University) uses MIMIC-III and MIMIC-IV datasets.
- Multi-Label Graph Foundation Models: “Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning” uses Humloc, PCG, Blogcatalog, PPI datasets.
- Urban Socioeconomic Prediction: “Harnessing the Synergy between LLM Agents and Knowledge Graphs…” by Zhilun Zhou et al. (Tsinghua University) uses WorldPop population data and Dianping review platform data alongside UrbanKG. Code: https://anonymous.4open.science/r/LLM_UrbanKG_Socioeconomic_Prediction-AB61/.
- Reactivity Prediction: “RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model…” uses Pistachio reactions dataset, Buchwald-Hartwig, Pd-catalyzed BH coupling, and proprietary C-N coupling/amide formation datasets.
- Medical Time Series: “Is Self-Pretraining really useful to improve diagnosis in medical Time Series?” by Omar Coser et al. (Università Campus Bio-Medico di Roma) uses CAMARGO 2021, Non-EEG Stress, and Gait in Parkinson’s Disease v1.0.0 datasets. Code: https://github.com/omarcoser/SPT-Medical-Time-Series.
- Skin Lesion Diagnosis: “Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning…” uses ISIC-2018 Challenge Task 3 and ISIC-2019 datasets.
- Visual Continuous Control: “Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control” uses DeepMind Control Suite (DMControl) and Atari100k benchmarks.
- Single-Cell Transcriptomics: “BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells” uses scBaseCount, STRING v12, CellBench-LS, Reactome pathway database.
- Multimodal Knowledge Graph Completion: “ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models…” uses FB15K-237 and DB15K multimodal datasets.
- Speaker Recognition: “Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning” by Zezhong Jin et al. (The Hong Kong Polytechnic University) uses VoxCeleb2, VoxCeleb1, VoxSRC 2021 validation set, MUSAN, RIR datasets. Code: https://github.com/modelscope/3D-Speaker.
- 6G Vision Communication: “Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications” provides a survey of various models and benchmarks for deep media compression and generative communication.
- Video-to-Audio Generation: “Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation” uses VGGSound and AudioSet datasets.
- Federated Outlier Detection: “Attention, Anomalies! Handling Attention Layers in Unsupervised Federated Outlier Detection” uses KDDCUP10, NSL KDD, PAMAP2, LEAF datasets. Apricot library: https://github.com/bknie1/apricot.
- Node-Level Graph SSL: “NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning” uses 5 node classification benchmarks. Code: https://github.com/OliverZ-dot/Node-Jepa.
- Molecular GNNs for BBBP: “Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs…” uses the BBBP dataset from B3DB.
- SST Downscaling: “Transferable Dual-Stream Representations for Mesoscale-Preserving Sea Surface Temperature Downscaling” uses ERA5 atmospheric reanalysis and MUR SST product. Code: https://anonymous.4open.science/r/EddyFlow.
- Traffic Forecasting: “Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing” uses a China Telecom Shanghai cellular network dataset. Dataset: http://sguangwang.com/TelecomDataset.html.
- Federated Prognostics: “Robust and Personalized Federated Learning for Aircraft-Engine Prognostics…” uses NASA C-MAPSS turbofan benchmark. Code: https://github.com/robusfl/robust-personalized-fl-turbofan.
- Tactile Sensing: “Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays” uses the STAG benchmark. HuggingFace models: huggingface.co/EximiusLabs, session-memory layer: pip install engram-robomem.
- Multimodal Mediation Analysis: “AI-driven Multimodal Representation Learning for Latent Mediation Structure Discovery…” uses All of Us Research Program data. Code: https://github.com/congca/Causal-AI-framework-of-Environmental-and-Psychosocial-Determinants-of-Cognitive-Difficulties.
- Multimodal Emotion Learning: “C2MOE: Consistency and Complementarity-guided Mixture of Experts…” uses CMU-MOSI and CMU-MOSEI datasets.
- Multi-View Stereo: “LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation” uses ScanNetv2, ScanNet++, 7-Scenes, LIBERO, RoboTwin 2.0 datasets and leverages Depth Anything v2, StableNormal, MobileSAM, SigLIP, DINOv2, VGGT.
- Multimodal Sentiment Analysis: “Rethinking Modality Reliability in Multimodal Sentiment Analysis…” uses CMU-MOSI, CMU-MOSEI, CH-SIMS datasets.
- Focal Cortical Dysplasia Segmentation: “CRIL-U-Net: Compact Ratio-Interaction Learning for Focal Cortical Dysplasia Segmentation…” uses the FCD dataset from Schuch et al. (2023).
- Diffusion Transformers: “DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers” uses DINOv2-B, MAE, MoCov3 and ImageNet 256×256/512×512, PDB, GEOM-Drug, QM9 datasets.
- Flexible Job Shop Scheduling: “PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning…” leverages liquid neural network dynamics.
- Software Fault Localization: “HyperFL: Query-Adaptive Representation Learning for Software Fault Localization” uses SWE-bench Lite benchmark.
- Molecular Representation Learning: “Learning Molecular Representations from Cellular Phenotypes with Structure Preservation” by Xuan Lin et al. (Xiangtan University, Hunan University) uses JUMP-CP dataset and Broad Drug Repurposing Hub. Code: https://github.com/JacklinGroup/Phenmol.
Impact & The Road Ahead
These papers collectively paint a picture of a field relentlessly pushing for more robust, efficient, and interpretable representations across diverse domains. The impact is far-reaching:
- Medical AI stands to gain immensely from semantic-aware multimodal pre-training for tabular data, disentangled skin lesion representations, motion artifact reduction in MRI, and biologically informed multi-omics foundation models. The potential for more accurate diagnostics, personalized treatments, and accelerated drug discovery is immense.
- Robotics and Autonomous Systems will benefit from improved 3D scene understanding, teachable 2D-3D matching, and robust tactile perception from low-cost sensors. This directly translates to more capable and safer intelligent agents.
- Scientific Discovery in chemistry, biology, and climate science is being accelerated by transformation-aware reaction models, graph-connected gene block predictions for single-cell analysis, and physics-informed models for sea surface temperature downscaling. These tools empower researchers to unlock new insights from complex data.
- Software Engineering and Security are seeing breakthroughs in binary code analysis with instruction-level alignment and query-adaptive fault localization, making software more secure and development more efficient.
- Resource-Constrained AI: The emphasis on lightweight architectures, efficient pretraining, and incremental learning (as seen in federated settings, small-data time series, and job shop scheduling) means that advanced AI capabilities can be deployed in environments previously deemed infeasible, from IoT devices to developing nations.
The road ahead involves further exploration of domain-specific inductive biases, pushing the boundaries of unsupervised disentanglement, and developing more sophisticated ways to integrate multimodal information. The synergy between theoretical foundations (like Landau theory for invariant learning and functional orthogonality for disentanglement) and practical innovations (like specialized attention mechanisms and adaptive fusion strategies) will continue to drive the field forward. As AI systems become more ubiquitous, the quality of their underlying representations will be paramount to their reliability, fairness, and overall societal benefit. The future of representation learning is not just about what models learn, but how they learn it, with an increasing focus on efficiency, robustness, and biological or physical plausibility.
Share this content:
Discover more from SciPapermill
Subscribe to get the latest posts sent to your email.
Post Comment