The year/Topics/Interpretability and analysis

Research collection

Interpretability and analysis

Understanding why models behave as they do: mechanistic interpretability, sparse autoencoders, probing, scaling and empirical laws, emergent phenomena, and hallucination analyses.

Papers
23
Research labs
3
Official code
11

123 of 23 papers in this collection

01

Independent research

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

The paper investigates how text-to-image diffusion transformers (DiTs) incorporate text semantics, focusing on chat-template tokens introduced by LLM-based text encoders. Using a causal interpretability framework on Qwen-Image models, the authors find that template tokens, despite carrying little prompt-specific information, become dominant attention sinks…

Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, et al.
Published
Jul 2026
Citations
0
Code
9 stars
02

Independent research

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

The paper identifies a cause of LLMs' suboptimal zero-shot text embedding performance: text embeddings align with high-frequency but uninformative tokens when projected onto the vocabulary space. Using Logit Lens and Logit Spectroscopy, the authors discover an 'edge spectrum' subspace in the unembedding matrix that encodes these frequent tokens. They…

Songhao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui, et al.
Published
Jun 2026
Citations
0
Code
25 stars
03

Independent research

On the Geometry of On-Policy Distillation

This paper analyzes the parameter-space geometry of on-policy distillation (OPD) for large language models, comparing it with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). Using diagnostics like update sparsity, subspace rotation, spectral drift, and update localization, the authors find that OPD occupies a…

Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, et al.
Published
Jun 2026
Citations
2
Code
Not linked
04

arXiv.org

From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

BrainCause is an automated framework for causally discovering and validating visual concept representations in the human brain using fMRI. It addresses the limitation of activation-based methods, which often identify false positives driven by correlated visual or semantic cues. BrainCause constructs targeted stimulus sets with positive images,…

Yuval Golbari, Navve Wasserman, Matias Cosarinsky, Roman Beliy, et al.
Published
May 2026
Citations
0
Code
Not linked
05

Independent research

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

This paper introduces an axiomatic evaluation framework for latent thought representations in LLMs, defining four functional axioms: Causality, Minimality, Separability, and Stability. Each axiom is quantified by a metric computed directly on the representation, independent of downstream task accuracy. The authors audit five open-weight LLMs (Llama-3.1 8B,…

Fahd Seddik, Fatemeh Fard
Published
May 2026
Citations
0
Code
7 stars
06

arXiv.org

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

This survey is the first comprehensive review of Attention Sink (AS) in Transformers, a phenomenon where disproportionate attention is focused on a small set of uninformative tokens. The authors synthesize over 210 studies, organizing the field into three key dimensions: Fundamental Utilization (e.g., Sink Token Preservation, Attention Redistribution,…

Zunhai Su, Hengyuan Zhang, Wei Wu, Yifan Zhang, et al.
Published
Apr 2026
Citations
3
Code
138 stars
07

Google DeepMind

Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

The paper investigates why reasoning improves parametric knowledge recall in LLMs for simple, single-hop factual questions. Using hybrid models (Gemini-2.5-Flash, Gemini-2.5-Pro, Qwen3-32B) on SimpleQA-Verified and EntityQuestions, the authors find that enabling reasoning substantially expands the model's capability boundary, as measured by pass@k, with…

Zorik Gekhman, Roee Aharoni, Eran Ofek, Mor Geva, et al.
Published
Mar 2026
Citations
9
Code
Not linked
08

OpenAI

Reasoning Models Struggle to Control their Chains of Thought

The paper introduces CoT-Control, an evaluation suite with 14,076 instances, to measure CoT controllability—the ability of reasoning models to follow instructions that constrain their chain of thought (CoT). Across 13 frontier models, CoT controllability is significantly lower than output controllability (e.g., Claude Sonnet 4.5: 2.7% vs 61.9%).…

Chen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He, et al.
Published
Mar 2026
Citations
11
Code
52 stars
09

arXiv.org

Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?

This paper evaluates whether Sparse Autoencoders (SAEs) recover meaningful features from neural networks. In synthetic experiments with known ground-truth features, SAEs achieved 71% explained variance but recovered only 9% of true features, showing a disconnect between reconstruction fidelity and feature recovery. On real LLM activations, the authors…

Anton Korznikov, Andrey Galichin, Alexey Dontsov, Oleg Rogov, et al.
Published
Feb 2026
Citations
7
Code
Not linked
10

Anthropic

Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs

This paper presents the first application of crosscoders to cross-architecture model diffing, introducing Dedicated Feature Crosscoders (DFCs) to better isolate model-exclusive features. DFCs partition the feature space into model-exclusive and shared sets, overcoming the standard crosscoder's prior toward shared features. In synthetic toy models, DFCs…

Thomas Jiralerspong, Trenton Bricken
Published
Feb 2026
Citations
9
Code
Not linked
11

Conference of the European Chapter of the Association for Computational Linguistics

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders

This paper introduces AudioSAE, the first large-scale application of Sparse Autoencoders (SAEs) to audio models, training them on all encoder layers of Whisper and HuBERT. The authors evaluate feature stability, interpretability, and practical utility. Over 50% of features remain consistent across random seeds, and reconstruction quality is preserved. SAE…

Georgii Aparin, Tasnima Sadekova, Alexey Rukhovich, Assel Yermekova, et al.
Published
Feb 2026
Citations
5
Code
15 stars
12

Independent research

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

The paper investigates the latent planning horizon of Large Language Models (LLMs) during Chain-of-Thought (CoT) reasoning. The authors introduce Tele-Lens, a probing method using low-rank adapters to predict teleological information (subsequent tokens, final answers, reasoning length) from hidden states across 12 diverse tasks. Empirical results reveal…

Liyan Xu, Mo Yu, Fandong Meng, Jie Zhou
Published
Feb 2026
Citations
1
Code
7 stars
13

Anthropic

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

This paper presents the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude.ai conversations using a privacy-preserving approach. The authors define situational disempowerment as occurring when interactions risk leading users to distorted perceptions of reality,…

Mrinank Sharma, Miles McCain, Raymond Douglas, David Duvenaud
Published
Jan 2026
Citations
19
Code
Not linked
14

Anthropic

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks

This paper introduces enhanced Constitutional Classifiers, a production-grade defense system against universal jailbreaks for large language models. The authors identify vulnerabilities in previous-generation defenses, such as reconstruction and output obfuscation attacks, and address them with exchange classifiers that evaluate outputs in full…

Hoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic, et al.
Published
Jan 2026
Citations
28
Code
Not linked
15

arXiv.org

Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits

The paper introduces Gnosis, a lightweight self-awareness mechanism that enables frozen large language models (LLMs) to predict their own failures by decoding signals from internal hidden states and attention patterns during inference. Gnosis compresses these internal traces into fixed-budget descriptors using a dual-stream architecture (hidden-state and…

Amirhosein Ghasemabadi, Di Niu
Published
Dec 2025
Citations
10
Code
46 stars
16

Conference of the European Chapter of the Association for Computational Linguistics

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

This paper provides the first comprehensive study of intrinsic dimension (ID) in text representations, grounding it in interpretable properties through cross-encoder analysis, linguistic features, and sparse autoencoders (SAEs). The authors establish three key findings: (1) ID is complementary to entropy-based metrics, as after controlling for length, the…

Vladislav Pedashenko, Laida Kushnareva, Yana Khassan Nibal, Eduard Tulchinskii, et al.
Published
Nov 2025
Citations
1
Code
Not linked
17

arXiv.org

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

AraLingBench is a fully human-annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). It consists of 150 expert-designed multiple-choice questions across five categories: grammar, morphology, spelling, reading comprehension, and syntax. The benchmark was constructed by five Arabic linguistics experts from the…

Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Sina Mukalled, Nadine Rizk, et al.
Published
Nov 2025
Citations
1
Code
9 stars
18

OpenAI

Weight-sparse transformers have interpretable circuits

The paper introduces weight-sparse transformers, where most weights are zero, to improve mechanistic interpretability. By constraining the L0 norm, models learn disentangled circuits for tasks, which are pruned to isolate minimal circuits. These circuits are compact and contain neurons and residual channels corresponding to natural concepts, with…

Leo Gao, Achyuta Rajaram, Jacob Coxon, Soham V. Govande, et al.
Published
Nov 2025
Citations
37
Code
Not linked
19

arXiv.org

Language Models are Injective and Hence Invertible

This paper proves that decoder-only Transformer language models are almost surely injective: distinct input prompts map to distinct last-token hidden representations, both at initialization and after any finite number of gradient descent steps. The authors establish this by showing the model is a real-analytic function of its parameters, so collisions can…

Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, et al.
Published
Oct 2025
Citations
31
Code
26 stars
20

Conference on Empirical Methods in Natural Language Processing

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

The paper introduces PsiloQA, a large-scale multilingual dataset for span-level hallucination detection in LLMs, covering 14 languages. It is built via an automated pipeline: generating QA pairs from Wikipedia using GPT-4o, eliciting answers from diverse LLMs without context, annotating hallucinated spans with GPT-4o, and filtering low-quality samples. The…

Elisei Rykov, Kseniia Petrushina, Maksim Savkin, Valerii Olisov, et al.
Published
Oct 2025
Citations
10
Code
Not linked
21

OpenAI

Why Language Models Hallucinate

The paper argues that language model hallucinations arise from statistical pressures during pretraining and persist due to misaligned evaluation metrics. The authors formalize hallucinations as errors in binary classification, showing that even with error-free training data, the cross-entropy objective leads to errors. They introduce the Is-It-Valid (IIV)…

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang
Published
Sep 2025
Citations
282
Code
Not linked
22

AAAI Conference on Artificial Intelligence

Beyond Transcription: Mechanistic Interpretability in ASR

This paper adapts interpretability methods from LLMs—logit lens, linear probing, and activation patching—to analyze the internal mechanisms of ASR models, specifically Whisper-large-v3 and Qwen2-Audio. The authors find that acoustic and semantic attributes (e.g., speaker gender, noise, accent) are linearly decodable from encoder layers, with peak…

Neta Glazer, Yael Segal-Feldman, Hilit Segev, Aviv Shamsian, et al.
Published
Aug 2025
Citations
11
Code
Not linked
23

arXiv.org

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models

The paper introduces FinCDM, the first cognitive diagnosis evaluation framework for financial large language models (LLMs), moving beyond aggregate scores to assess knowledge-skill level proficiency. It constructs CPA-KQA, a dataset of 210 expert-annotated questions derived from the CPA exam, covering 70 financial concepts, with high inter-annotator…

Ziyan Kuang, Feiyu Zhu, Maowei Jiang, Yanzhao Lai, et al.
Published
Aug 2025
Citations
4
Code
3 stars