Loading Now

Interpretability Unleashed: Unpacking the Black Box from Neurons to Ecosystems

Latest 100 papers on interpretability: Oct. 10, 2026

The quest for interpretability in AI and Machine Learning continues to be a central, multifaceted challenge. As models grow in complexity and autonomy, understanding their internal workings, decision-making processes, and potential failure modes becomes paramount for trust, safety, and effective deployment. Recent research showcases a vibrant landscape of breakthroughs, pushing the boundaries from theoretical foundations to practical, real-world applications. This digest dives into these advancements, revealing how researchers are shedding light on the “black box” across diverse domains.

The Big Idea(s) & Core Innovations

One overarching theme is the move beyond purely post-hoc explanations to intrinsically interpretable or mechanistically verifiable models. For instance, The Polytopal Neural Network by A. Emilie J. Wedenborg, Anders V. Nørskov, and their team from the Technical University of Denmark [https://arxiv.org/pdf/2610.12004] introduces PNNs, which constrain latent representations to lie on polytopes. This provides integrated interpretability: the same coordinates used for explanation are directly used for computation, ensuring faithfulness. Similarly, HGNAN (Hypergraph Neural Additive Network) by Shihan Feng, Xin Zheng, et al. from UNC Chapel Hill [https://arxiv.org/pdf/2610.07458] extends neural additive models to hypergraphs, replacing entangled message-passing with a decomposable additive structure for direct attribution of predictions to features and structural components.

Another significant thrust is the application of mechanistic interpretability to understand and steer complex foundation models. FearCaut-Qwen by Xiaoshan Zhou from the University of Sydney [https://arxiv.org/pdf/2610.11986] demonstrates how manipulating a fear/threat neural circuit in a Vision-Language Model can shift its hazard assessment decision criterion at inference time, improving recall without retraining. Ilya Lasy et al. from TU Wien [https://arxiv.org/pdf/2610.11775], in their paper RouterInterp, challenge the “domain specialization hypothesis” in Mixture of Experts (MoE) models, proposing the Superposed Specialisation Hypothesis. They use Sparse Autoencoders (SAEs) to identify that experts specialize in disjoint collections of unrelated micro-domains, and their RouterInterp method can generate unified natural language explanations with significantly higher accuracy than prior approaches. Building on this, SafeEvo by Miao Yu et al. from the University of Hong Kong [https://arxiv.org/pdf/2610.09600] attributes LLM safety alignment to sparse refusal circuits, even in pretrained models, and introduces Safety Circuit Alignment (SCA) to confine safety updates to these circuits, leading to better alignment and reduced over-refusal. Further pushing this concept of internal model manipulation, Muhammad Atif Butt et al. in Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models [https://arxiv.org/pdf/2610.06945] introduce KANSteer, using Kolmogorov-Arnold Networks to model visual concept traversal as curved trajectories in activation space, allowing more faithful steering than linear methods.

The drive for interpretability is also deeply integrated into domain-specific AI systems. Sera by Jiawei Li et al. from Singapore Institute of Technology [https://arxiv.org/pdf/2610.11567] integrates structured text-based degradation semantics with numerical battery time series for interpretable State of Health (SoH) forecasting, reducing prediction error by up to 37.3%. For medical imaging, Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention by Samrajya Thapa et al. from Iowa State University and Mayo Clinic [https://arxiv.org/pdf/2610.09031] proposes a plug-and-play framework for concept-based interpretation and refinement, allowing clinicians to debug models and isolate causal from spurious correlations. Similarly, X-OPM by Jingbo Jiang et al. from HKUST [https://arxiv.org/pdf/2610.08502] creates a white-box power modeling framework for digital on-chip power meters, leveraging RTL design principles for strictly explainable features and improved robustness. In a similar vein, DISSOLVR by Vansh Ramani et al. from IIT Delhi [https://arxiv.org/pdf/2610.02574] achieves state-of-the-art solubility prediction using interpretable Gradient Boosted Decision Trees and LLM-assisted explanations, challenging the necessity of deep learning for top performance. EpiWorld by Zeeshan Memon et al. from Emory University [https://arxiv.org/pdf/2610.02744] grounds LLM policy agents in epidemiological world models for counterfactual policy evaluation, generating interpretable, protocol-compliant interventions that reduce hospitalizations.

For symbolic reasoning, SED-MCTS by Yunpeng Gong et al. from Xiamen University [https://arxiv.org/pdf/2610.12003] distills structural experience from evaluated PDE candidate expressions to guide symbolic solution search, providing an interpretable approach to PDE discovery. Towards the Automatic Synthesis of Interpretable Chess Tactics by Abhijeet Krishnan and Chris Martens from NC State University [https://arxiv.org/pdf/2610.07640] uses Inductive Logic Programming to learn symbolic chess tactics expressed as first-order logic rules, generating human-interpretable move suggestions.

Crucially, several papers tackle the reliability and rigorous evaluation of interpretability methods themselves. Pranjal Garg in How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing [https://arxiv.org/pdf/2610.08544] introduces a floor-ceiling-headroom framework to normalize probe scores, arguing that raw scores lack fixed meaning and can be inflated by input leakage. Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability by Chuqin Geng et al. from the University of Toronto and McGill University [https://arxiv.org/pdf/2610.02098] reveals a fundamental flaw: faithfulness metrics can systematically misrank equally-sized circuits, preferring less behaviorally accurate ones due to context distortion. This work underscores the need for more robust evaluation protocols. How to train your model organism by Xilin Wang, David Bau, and Byron C. Wallace from Northeastern University [https://arxiv.org/pdf/2610.10203] proposes a multi-objective validation framework for LLM “model organisms” used in interpretability research, ensuring they preserve general capabilities and naturalness alongside target behavior, to make interpretability conclusions more reliable.

Under the Hood: Models, Datasets, & Benchmarks

Recent advancements often hinge on specialized models, novel datasets, and robust benchmarks:

  • Polytopal Neural Networks (PNNs): Evaluated on image classification tasks using datasets like MNIST, FashionMNIST, CIFAR-10, SVHN, EuroSAT, MedMNIST v2, and ImageNet variants. Leverages ConvNeXt-T backbone.
  • EDMD-kDL: Kernel-based autoencoder. Evaluated on NOAA sea surface temperature reanalysis, pendulum, fluid flow, and video data. Benchmarked against ANN-based Koopman autoencoders and SINDy-SHRED. No specific code repository provided, but baselines are available at github.com/erichson/koopmanAE and github.com/pyshred-dev/pyshred.
  • FearCaut-Qwen: A modified Qwen2.5-VL-7B-Instruct model with affective steering. Uses the novel SeisMLLM-1K dataset (1,306 building cases, 3,421 images from 17 earthquakes) and FindingEmo for emotion recognition. Implemented with PyTorch 2.14.0 and Transformers 5.16.1.
  • RouterInterp: Employs Sparse Autoencoders (SAEs) on gpt-oss-20b and OLMoE-1B-7B. Code is available at https://github.com/ilyalasy/routerinterp along with the Delphi library.
  • SED-MCTS: A Monte Carlo Tree Search framework. Paper notes resources at https://arxiv.org/pdf/2610.12003 for reference.
  • Sera: Three-branch architecture combining temporal modeling (PatchTST) with rule-based and LLM-based (via LLMs not specified but for interpretation) semantic representations. Evaluated on the MIT Battery Degradation Dataset. Resources at https://arxiv.org/pdf/2610.11567.
  • MARI (Memory-Augmented Recommendation with Interpretability): Leverages Qwen3 series models (Qwen3-4B to Qwen3-235B-A22B) and DeepSeek-R1 for distillation. Uses Qwen3-Embedding-0.6B for vectorization. No code provided.
  • Kolmogorov-Arnold Networks (KANs): Compared against MLPs. Evaluated on Feynman physics benchmark and Gymnasium RL environments (Acrobot, CartPole, MountainCar, Pendulum). Code for KAN training: https://github.com/DerKevinRiehl/neurips26_kan_training. SW-KAN [https://arxiv.org/abs/2610.00050] and Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks [https://arxiv.org/pdf/2610.00985] further advance KANs with specialized polynomial activations and scaling law analysis. Code for SW-KAN: https://github.com/amirhoseinazarpour/SW-KAN.
  • Strategic Governance of AI Models in Earth Science: White paper leveraging models like Aurora and Prithvi WxC, and datasets like WeatherBench 2. Advocates for new open AI-ready evaluation datasets and shared reporting standards. Resources at https://arxiv.org/pdf/2610.10560.
  • Temporally Interpretable Differentiable Decision Trees (DDTs): Uses action chunking and Information-Theoretic Tree Restructuring (ITTR). Tested on Leurent’s Lane Keeping and Gymnasium environments (Inverted Pendulum, Lunar Lander). Code: https://github.com/ei5uke/temp-interp.
  • LLMT: Distills Qwen2.5-72B-Instruct into decision trees. Evaluated on 11 UCI ML Repository datasets (e.g., Diabetes, Spambase, Nursery) and an Ecom dataset. Code: https://github.com/yueqiu0/LLMTree.
  • Model Organisms for Interpretability Research: New clinical model organism suite for demographic biases. Prefers DPO over SFT and uses model merging. HuggingFace repo: https://huggingface.co/multi-objective-mo. Code: https://github.com/Rice-wxl/multi_objective_mo.
  • VLA Language Grounding Study: Mechanistic interpretability on π0.5 and GR00T N1.7. Uses LeRobot’s Libero_10_image dataset. Code for NNsight framework (to be released).
  • Explainable ML for Textile Pressure-Based Postural Screening: XGBoost with engineered features. Uses the Smart-Sleeve dataset: https://github.com/xghgithub/Smart-Sleeve-Dataset.
  • Extreme Binary Classification: Uses Extreme Value Theory with threshold adaptation. Evaluated on Credit Card Fraud, Brazilian Corruption, Breast Cancer Wisconsin, and ScreenX datasets. Utilizes imbalanced-learn, MAPIE, nproc, XGBoost, and scikit-learn.
  • Identifiability of a dissipative knowledge-dynamics model: Dissipative ODE model. Uses ASSISTments-2009, ASSISTments-2015, ASSISTments-2017, and Junyi-2015 datasets. Code: https://github.com/armankostanian/cognitive-pinn.
  • Fully Interpretable Minimal Transformers: Minimal 2D transformer. Code for ‘plus-last-even’ task: https://github.com/Raneem-mahajne/creating_transformer/tree/main/plus_last_even.
  • U-SPACE: Training-free framework for uncertainty quantification in LLMs. Uses reasoning models and benchmarks like J-LENS. Code: https://github.com/s2labres/U-Space.
  • CXPMRG-Bench & MambaXray-PRB: New benchmark and framework for X-ray report generation using Mamba vision encoder and LLM decoder. Evaluated on IU X-ray, MIMIC-CXR, and CheXpert Plus datasets. Code: https://github.com/Event-AHU/Medical Image Analysis.
  • PerSpectron: Hardware-based detection system using perceptrons. Evaluated with gem5 simulator, FANN C library, scikit-learn, and SPEC CPU 2006 benchmarks.
  • Sera: Framework for battery health forecasting. Utilizes MIT Battery Degradation Dataset. Resources at https://arxiv.org/pdf/2610.11567.
  • MoTIF-X: Motif-centered token integration framework for molecular representation learning. Evaluated on OpenADMET ExpansionRx, GEOM-Drugs, MoleculeNet, BindingDB, ChEMBL v35, BIOSNAP, DAVIS datasets. No code provided.
  • EvoRiskBench: Execution-grounded benchmark for workspace agents, using EP-Path-EF framework. Evaluates DeepSeek-V4-Pro-0813 with Codex. Resources at https://arxiv.org/pdf/2610.03153.
  • Weights Oracles: LLMs reading raw neural network weights. Works with small transformers. Resources at https://arxiv.org/pdf/2610.07334.
  • RL-PaO: Reinforcement Learning framework for prediction as action. Evaluated on day-ahead energy scheduling using Japan Meteorological Agency dataset. No code provided.
  • Seeing Time (WAVE): Visual-temporal MTS clustering framework. Evaluated on 10 UEA multivariate time series datasets. Code: https://github.com/Zheng-Zhu1/WAVE.
  • Sociality Anchors: Trajectory prediction framework. Evaluated on ETH-UCY and Stanford Drone Dataset (SDD). Code: https://github.com/LivepoolQ/Socialality.
  • RESCUE: Framework for repairing LLM errors to sparse circuits. Uses Qwen3-8B and Llama-3.1-8B-Instruct. Evaluated on GSM8K, MATH-500, MedMCQA, and other benchmarks. Code: https://github.com/chuanpupig/RESCUE.
  • NeuroDyn-EEG: EEG pretraining framework based on neural dynamics. Uses an extended Jansen-Rit neural mass model. Evaluated on clinical benchmarks for Alzheimer’s, Parkinson’s, and MDD. Code: https://github.com/Gnosis-Neurodynamics/NeuroDyn-EEG.
  • LAURA: Knowledge distillation for legal NLP. Uses GPT-4o teacher to Flan-T5 student. Evaluated on CUAD and a contract ambiguity dataset. No code provided.
  • Neural Structural Reasoner (NSR): Brain-inspired architecture for knowledge graph reasoning. Evaluated on Nations, Kinship, YAGO3-10, FB15k-237, Countries S3, WN18RR datasets. No code provided.
  • Factorized Scheduling Principle (FSP): Interpretable scheduling rules via structured additive functions. Code for implementation not explicitly provided but conceptualized as RL framework.
  • DraftTrace: Writing environment for AI-integrated writing analytics. No code provided.
  • Emergent Tonal Structure: Skip-gram chord embeddings. Uses BPS-FH corpus (Beethoven) and Isophonics (Beatles). Code: https://github.com/MaraalE/chord-tonality.
  • Receptive-field-constrained stimulus optimization (RF-DiVE/RF-GO): MEI generation for visual cortex. Uses Natural Scenes Dataset (NSD) and LAION-fMRI. No code provided.
  • Active Budget Can Kill Sensitivity: Diagnosing TopK Sparse Autoencoder Reliability. Uses GPT-2 small, Qwen2.5-1.5B-Instruct, Gemma-2-9b. No code provided.
  • CHOQOLATE: Concept Bottleneck Models with Choquet Integrals. Uses CLIP ViT-L/14@336px. Evaluated on Cats/Dogs/Cars, MonumAI, COCO, CUB-200. No code provided.
  • Weights Read and Write Features (ASPD): Activation-Supported Parameter Decomposition. Uses Qwen-3-8B, GPT-2 small, Gemma-2-2B. Code: https://github.com/tue147/weights-read-write.
  • When Models Don’t Manipulate Manifolds: Geometry of number comparison in Qwen2.5-7B-Instruct. Code: https://github.com/Sai-Sumedh/comparison-geometry.
  • RL-PaO: Prediction as Action in Decision Making under Uncertainty. Uses data from the Japan Meteorological Agency. No code provided.
  • REVEAL: Self-evolving framework for AI-generated video detection using VLMs. New VidForensic benchmark (1.4k+ videos). Uses Qwen-VL-Max, Gemini-1.5-pro, GPT-4o, Llava-OV-7B.
  • AttSVD: Training-free KV cache compression. Evaluated on LongBench and an agentic benchmark. Core code released as supplementary material.
  • Probabilistic Truly Unordered Rule Sets (TURS): Rule-set model with MDL-based learning.
  • Fast, Interpretable, and Deterministic Time Series Classification With a Bag-of-Receptive-Fields (BORF): Evaluated on 150+ UEA/UCR datasets. Code: https://github.com/fspinna/borf.
  • Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers: Investigates DINOv2. Accompanying artifact with trained SAE checkpoints.
  • Certified Mechanistic Edits: Formal verification for neural network edits. Resources at https://arxiv.org/pdf/2610.03502.
  • From Patching to Pruning Visual Computation in Vision Language Models (P2P): Training-free framework for VLM computation bypass. Uses Qwen2.5-VL and LLaVA families. No code provided.
  • X-ray Report Generation on CheXpert Plus Dataset (MambaXray-PRB): Framework with Mamba vision encoder. Code: https://github.com/Event-AHU/Medical Image Analysis.
  • Vectorized Dynamic Histograms for Sparse Oblique Forests: Optimizes Yggdrasil Decision Forests (YDF) with AVX2/AVX-512. Code at anonymized link and within YDF library.
  • NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning. Uses Sparse Autoencoders on Qwen2.5-VL-7B and LLaVA-1.6-7B. No code provided.
  • Which Attention Heads are Like the Human Head?: Investigates brain-AI alignment in LLMs using EEG. No code provided.
  • Beyond Linear Concepts: Non-Linear Concept Manifolds in LLMs. Uses NLMCD from computer vision. Code: https://anonymous.4open.science/r/NLMCD-NLP-C5E7.
  • End-to-End Learning vs. Modular Architectures: Comparative analysis for autonomous driving. Uses nuScenes, CARLA, DARPA Urban Challenge.
  • Iterative Policy Structure Refinement: LLM-guided policy refinement for RL. No code provided.
  • Have an LLM Write Your Anomaly Detector: LLM-driven autonomous research loop for time series anomaly detection. Code: https://anonymous.4open.science/r/TS-AD_submission-7E42/README.md.
  • Detect, Explain, Interpret (SHAD): Time series anomaly detection benchmark. Code: https://github.com/scality/shad/blob/main/README.md.
  • Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks: Code for BSRBF-KAN: https://github.com/hoangthangta/BSRBF-KAN and Faster-KAN: https://github.com/AthanasiosDelis/faster-kan/.
  • Geometric Similarity in VLM Low-Level Vision Representations: GeoSim framework for VLM representation analysis. Uses multiple VLM architectures (Emu3-Chat, Anole-7B, Janus-Pro-7B, InternVL3.5-8B, Qwen3-VL-8B, Qwen-Image-Edit, Emu3.5-Image). No code provided.
  • Interpreting Reasoning of Large Language Models via Partial Information Decomposition (SLIDER): Uses PID for LLM reasoning analysis. Code not explicitly provided, but utilizes GSM8K and AIME25 datasets.
  • Adaptive Conformal Prediction for Image Regression (ACPNN): For ICF emulation. Uses BigFoot ICF simulation dataset. Code for GpGp R package (Guinness, 2019).
  • Interpretable Synthetic Medical Tabular Data Generation Using Fuzzy Cognitive Maps: Uses UCI Pima Indians Diabetes, South African Heart Disease, Statlog Heart Disease. No code provided.
  • Decoding the Disaster (GRDisaster): Multi-task geospatial reasoning with VLMs. Uses PhotoMappers benchmark. No code provided.
  • Legal text classification in Korean sexual offense cases: Uses KLUE-BERT, KPF-BERT, LBox-lcube, Llama-3.1-8B, Polyglot-ko. Code for transformers-interpret XAI framework.

Impact & The Road Ahead

The collective impact of this research is profound. From enhancing the safety and reliability of autonomous systems (as highlighted in End-to-End Learning vs. Modular Architectures [https://arxiv.org/pdf/2610.01746] and Strategic Governance of AI Models in Earth Science [https://arxiv.org/pdf/2610.10560]) to improving clinical decision-making with explainable AI (Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention [https://arxiv.org/pdf/2610.09031], CXPMRG-Bench [https://arxiv.org/pdf/2610.08813], NeuroDyn-EEG [https://arxiv.org/pdf/2609.36773], Towards Trustworthy AI for Glioma Diagnosis [https://arxiv.org/pdf/2609.39429]), interpretability is moving from a theoretical ideal to a practical necessity. The development of frameworks like RolloutFaith [https://arxiv.org/pdf/2609.36843] for auditing persistent internal interventions and Certified Mechanistic Edits [https://arxiv.org/pdf/2610.03502] for formal guarantees mark a significant step towards truly trustworthy AI.

Looking ahead, several exciting directions emerge. The convergence of physics-informed models with deep learning, exemplified by EDMD-kDL [https://arxiv.org/pdf/2610.12370] and Atoms to Processes [https://arxiv.org/pdf/2610.02014], promises more robust and generalizable solutions for scientific discovery and engineering. The ability to program or steer models through interpretable concepts (NeuronEye [https://arxiv.org/pdf/2609.38098], FearCaut-Qwen [https://arxiv.org/pdf/2610.11986], Beyond the Linear Representation Hypothesis [https://arxiv.org/pdf/2610.06945]) suggests a future where humans and AI collaborate more intimately, with humans providing high-level guidance and AI executing complex tasks transparently. The rise of Kolmogorov-Arnold Networks (Sample-Efficiency of Kolmogorov-Arnold Networks [https://arxiv.org/pdf/2610.10627], SW-KAN [https://arxiv.org/abs/2610.00050]) offering both interpretability and efficiency is particularly exciting for resource-constrained and safety-critical applications. Finally, the meta-interpretability research, such as How High Is 0.6? [https://arxiv.org/pdf/2610.08544] and Are We Recovering Mechanisms? [https://arxiv.org/pdf/2610.02098], is crucial for ensuring that our interpretability tools are themselves reliable. The journey to truly understand and control advanced AI is long, but these recent papers demonstrate incredible progress, paving the way for a more transparent, robust, and ultimately, more useful AI future.

Share this content:

mailbox@3x Interpretability Unleashed: Unpacking the Black Box from Neurons to Ecosystems
Hi there 👋

Get a roundup of the latest AI paper digests in a quick, clean weekly email.

Spread the love

Discover more from SciPapermill

Subscribe to get the latest posts sent to your email.

Post Comment

Discover more from SciPapermill

Subscribe now to keep reading and get access to the full archive.

Continue reading