The year/Topics/Safety and analysis

Topic area

Safety and analysis

Every collection across safety and analysis.

Papers
35
Research labs
3
Official code
15

135 of 35 papers in this topic area

01

Independent research

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

The paper investigates how text-to-image diffusion transformers (DiTs) incorporate text semantics, focusing on chat-template tokens introduced by LLM-based text encoders. Using a causal interpretability framework on Qwen-Image models, the authors find that template tokens, despite carrying little prompt-specific information, become dominant attention sinks…

Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, et al.
Published
Jul 2026
Citations
0
Code
9 stars
02

OpenAI

Predicting LLM Safety Before Release by Simulating Deployment

This paper introduces deployment simulation, a method for predicting LLM safety before release by resampling the next assistant response from de-identified production conversation prefixes using a candidate model. The authors evaluate this approach across GPT-5-series deployments, finding that it produces informative estimates of post-deployment…

Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, et al.
Published
Jul 2026
Citations
1
Code
Not linked
03

bioRxiv

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

MedMisBench is a benchmark introduced to measure the epistemic resilience of large language models (LLMs) in medical settings, defined as the ability to maintain correct medical judgment when misleading context is present. It contains 10,932 medical question items and 48,889 misleading context-option pairs, built from five source datasets covering medical…

Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, et al.
Published
Jun 2026
Citations
1
Code
9 stars
04

Independent research

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

The paper identifies a cause of LLMs' suboptimal zero-shot text embedding performance: text embeddings align with high-frequency but uninformative tokens when projected onto the vocabulary space. Using Logit Lens and Logit Spectroscopy, the authors discover an 'edge spectrum' subspace in the unembedding matrix that encodes these frequent tokens. They…

Songhao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui, et al.
Published
Jun 2026
Citations
0
Code
25 stars
05

Independent research

On the Geometry of On-Policy Distillation

This paper analyzes the parameter-space geometry of on-policy distillation (OPD) for large language models, comparing it with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). Using diagnostics like update sparsity, subspace rotation, spectral drift, and update localization, the authors find that OPD occupies a…

Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, et al.
Published
Jun 2026
Citations
2
Code
Not linked
06

Independent research

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

AgentDoG 1.5, developed by Shanghai Artificial Intelligence Laboratory, is a lightweight and scalable framework for AI agent safety and security. It updates the three-dimensional safety taxonomy (risk source, failure mode, real-world harm) to cover new risks from Codex and OpenClaw execution scenarios, extending the ATBench benchmark family with…

Dongrui Liu, Yu Li, Zhonghao Yang, Peng Wang, et al.
Published
May 2026
Citations
2
Code
Not linked
07

arXiv.org

From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

BrainCause is an automated framework for causally discovering and validating visual concept representations in the human brain using fMRI. It addresses the limitation of activation-based methods, which often identify false positives driven by correlated visual or semantic cues. BrainCause constructs targeted stimulus sets with positive images,…

Yuval Golbari, Navve Wasserman, Matias Cosarinsky, Roman Beliy, et al.
Published
May 2026
Citations
0
Code
Not linked
08

arXiv.org

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

This survey provides a comprehensive analysis of Large Audio Language Models (LALMs), focusing on their generalization, trustworthiness, and future outlook. It examines the endogenous mechanisms of LALMs, including architectural foundations, representational paradigms, training and alignment strategies, and emergent reasoning mechanisms. The survey…

Kaiwen Luo, Zhenhong Zhou, Leyan Wang, Liang Lin, et al.
Published
May 2026
Citations
4
Code
258 stars
09

Independent research

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

This paper introduces an axiomatic evaluation framework for latent thought representations in LLMs, defining four functional axioms: Causality, Minimality, Separability, and Stability. Each axiom is quantified by a metric computed directly on the representation, independent of downstream task accuracy. The authors audit five open-weight LLMs (Llama-3.1 8B,…

Fahd Seddik, Fatemeh Fard
Published
May 2026
Citations
0
Code
7 stars
10

arXiv.org

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

This survey is the first comprehensive review of Attention Sink (AS) in Transformers, a phenomenon where disproportionate attention is focused on a small set of uninformative tokens. The authors synthesize over 210 studies, organizing the field into three key dimensions: Fundamental Utilization (e.g., Sink Token Preservation, Attention Redistribution,…

Zunhai Su, Hengyuan Zhang, Wei Wu, Yifan Zhang, et al.
Published
Apr 2026
Citations
3
Code
138 stars
11

OpenAI

IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

The paper introduces IH-Challenge, a reinforcement learning (RL) training dataset designed to improve instruction hierarchy (IH) robustness in large language models (LLMs). IH defines how models prioritize system, developer, user, and tool instructions under conflict, which is key for defending against jailbreaks, system prompt extractions, and prompt…

Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu, Christopher A. Choquette-Choo, et al.
Published
Mar 2026
Citations
12
Code
Not linked
12

Google DeepMind

Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

The paper investigates why reasoning improves parametric knowledge recall in LLMs for simple, single-hop factual questions. Using hybrid models (Gemini-2.5-Flash, Gemini-2.5-Pro, Qwen3-32B) on SimpleQA-Verified and EntityQuestions, the authors find that enabling reasoning substantially expands the model's capability boundary, as measured by pass@k, with…

Zorik Gekhman, Roee Aharoni, Eran Ofek, Mor Geva, et al.
Published
Mar 2026
Citations
9
Code
Not linked
13

OpenAI

Reasoning Models Struggle to Control their Chains of Thought

The paper introduces CoT-Control, an evaluation suite with 14,076 instances, to measure CoT controllability—the ability of reasoning models to follow instructions that constrain their chain of thought (CoT). Across 13 frontier models, CoT controllability is significantly lower than output controllability (e.g., Claude Sonnet 4.5: 2.7% vs 61.9%).…

Chen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He, et al.
Published
Mar 2026
Citations
11
Code
52 stars
14

arXiv.org

Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?

This paper evaluates whether Sparse Autoencoders (SAEs) recover meaningful features from neural networks. In synthetic experiments with known ground-truth features, SAEs achieved 71% explained variance but recovered only 9% of true features, showing a disconnect between reconstruction fidelity and feature recovery. On real LLM activations, the authors…

Anton Korznikov, Andrey Galichin, Alexey Dontsov, Oleg Rogov, et al.
Published
Feb 2026
Citations
7
Code
Not linked
15

Anthropic

Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs

This paper presents the first application of crosscoders to cross-architecture model diffing, introducing Dedicated Feature Crosscoders (DFCs) to better isolate model-exclusive features. DFCs partition the feature space into model-exclusive and shared sets, overcoming the standard crosscoder's prior toward shared features. In synthetic toy models, DFCs…

Thomas Jiralerspong, Trenton Bricken
Published
Feb 2026
Citations
9
Code
Not linked
16

arXiv.org

The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies

The paper argues that a multi-agent system built from large language models cannot simultaneously achieve continuous self-evolution, complete isolation from external feedback, and safety invariance—a combination termed the self-evolution trilemma. Using an information-theoretic framework, safety is formalized as the KL divergence from an anthropic value…

Chenxu Wang, Chaozhuo Li, Songyang Liu, Zejian Chen, et al.
Published
Feb 2026
Citations
3
Code
Not linked
17

Conference of the European Chapter of the Association for Computational Linguistics

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders

This paper introduces AudioSAE, the first large-scale application of Sparse Autoencoders (SAEs) to audio models, training them on all encoder layers of Whisper and HuBERT. The authors evaluate feature stability, interpretability, and practical utility. Over 50% of features remain consistent across random seeds, and reconstruction quality is preserved. SAE…

Georgii Aparin, Tasnima Sadekova, Alexey Rukhovich, Assel Yermekova, et al.
Published
Feb 2026
Citations
5
Code
15 stars
18

Independent research

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

The paper investigates the latent planning horizon of Large Language Models (LLMs) during Chain-of-Thought (CoT) reasoning. The authors introduce Tele-Lens, a probing method using low-rank adapters to predict teleological information (subsequent tokens, final answers, reasoning length) from hidden states across 12 diverse tasks. Empirical results reveal…

Liyan Xu, Mo Yu, Fandong Meng, Jie Zhou
Published
Feb 2026
Citations
1
Code
7 stars
19

Anthropic

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

This paper presents the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude.ai conversations using a privacy-preserving approach. The authors define situational disempowerment as occurring when interactions risk leading users to distorted perceptions of reality,…

Mrinank Sharma, Miles McCain, Raymond Douglas, David Duvenaud
Published
Jan 2026
Citations
19
Code
Not linked
20

Independent research

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

AgentDoG is a diagnostic guardrail framework for AI agent safety and security, developed by the Shanghai Artificial Intelligence Laboratory. It addresses limitations in existing guardrails by introducing a unified three-dimensional safety taxonomy that categorizes agentic risks by source (where), failure mode (how), and real-world harm (what). Guided by…

Dongrui Liu, Qihan Ren, Chen Qian, Shuai Shao, et al.
Published
Jan 2026
Citations
24
Code
680 stars
21

Anthropic

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks

This paper introduces enhanced Constitutional Classifiers, a production-grade defense system against universal jailbreaks for large language models. The authors identify vulnerabilities in previous-generation defenses, such as reconstruction and output obfuscation attacks, and address them with exchange classifiers that evaluate outputs in full…

Hoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic, et al.
Published
Jan 2026
Citations
28
Code
Not linked
22

arXiv.org

Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits

The paper introduces Gnosis, a lightweight self-awareness mechanism that enables frozen large language models (LLMs) to predict their own failures by decoding signals from internal hidden states and attention patterns during inference. Gnosis compresses these internal traces into fixed-budget descriptors using a dual-stream architecture (hidden-state and…

Amirhosein Ghasemabadi, Di Niu
Published
Dec 2025
Citations
10
Code
46 stars
23

OpenAI

Training LLMs for Honesty via Confessions

The paper proposes a method to train LLMs to produce 'confessions'—self-reports of compliance with instructions and policies—to improve honesty. Confession training adds a system message after the model's answer, requesting a structured report enumerating objectives, compliance analysis, and uncertainties. The confession reward is based solely on honesty…

Manas Joglekar, Jeremy Chen, Gabriel Wu, Jason Yosinski, et al.
Published
Dec 2025
Citations
17
Code
Not linked
24

Conference of the European Chapter of the Association for Computational Linguistics

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

This paper provides the first comprehensive study of intrinsic dimension (ID) in text representations, grounding it in interpretable properties through cross-encoder analysis, linguistic features, and sparse autoencoders (SAEs). The authors establish three key findings: (1) ID is complementary to entropy-based metrics, as after controlling for length, the…

Vladislav Pedashenko, Laida Kushnareva, Yana Khassan Nibal, Eduard Tulchinskii, et al.
Published
Nov 2025
Citations
1
Code
Not linked
25

arXiv.org

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

AraLingBench is a fully human-annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). It consists of 150 expert-designed multiple-choice questions across five categories: grammar, morphology, spelling, reading comprehension, and syntax. The benchmark was constructed by five Arabic linguistics experts from the…

Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Sina Mukalled, Nadine Rizk, et al.
Published
Nov 2025
Citations
1
Code
9 stars
26

OpenAI

Weight-sparse transformers have interpretable circuits

The paper introduces weight-sparse transformers, where most weights are zero, to improve mechanistic interpretability. By constraining the L0 norm, models learn disentangled circuits for tasks, which are pruned to isolate minimal circuits. These circuits are compact and contain neurons and residual channels corresponding to natural concepts, with…

Leo Gao, Achyuta Rajaram, Jacob Coxon, Soham V. Govande, et al.
Published
Nov 2025
Citations
37
Code
Not linked
27

Annual Meeting of the Association for Computational Linguistics

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains

This paper investigates the ability of large language models (LLMs) to role-play morally ambiguous or villainous characters, hypothesizing that safety alignment conflicts with authentic antagonistic portrayal. The authors introduce the Moral RolePlay benchmark, a dataset with a four-level moral alignment scale (Moral Paragons, Flawed-but-Good, Egoists,…

Zihao Yi, Qingxuan Jiang, Ruotian Ma, Xingyu Chen, et al.
Published
Nov 2025
Citations
7
Code
Not linked
28

arXiv.org

Language Models are Injective and Hence Invertible

This paper proves that decoder-only Transformer language models are almost surely injective: distinct input prompts map to distinct last-token hidden representations, both at initialization and after any finite number of gradient descent steps. The authors establish this by showing the model is a real-analytic function of its parameters, so collisions can…

Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, et al.
Published
Oct 2025
Citations
31
Code
26 stars
29

Anthropic

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

This paper investigates whether poisoning attacks on large language models (LLMs) require a constant number of poisoned samples regardless of dataset size, rather than a fixed percentage. The authors conducted the largest pretraining poisoning experiments to date, training models from 600M to 13B parameters on Chinchilla-optimal datasets (6B to 260B…

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies, et al.
Published
Oct 2025
Citations
76
Code
Not linked
30

Conference on Empirical Methods in Natural Language Processing

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

The paper introduces PsiloQA, a large-scale multilingual dataset for span-level hallucination detection in LLMs, covering 14 languages. It is built via an automated pipeline: generating QA pairs from Wikipedia using GPT-4o, eliciting answers from diverse LLMs without context, annotating hallucinated spans with GPT-4o, and filtering low-quality samples. The…

Elisei Rykov, Kseniia Petrushina, Maksim Savkin, Valerii Olisov, et al.
Published
Oct 2025
Citations
10
Code
Not linked
31

arXiv.org

Large Reasoning Models Learn Better Alignment from Flawed Thinking

Large reasoning models (LRMs) generate chain-of-thought (CoT) before answering but are easily biased by flawed reasoning prefills, leading to unsafe or overrefused outputs. The paper introduces RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a reinforcement learning (RL) post-training method that trains models to override flawed reasoning…

ShengYun Peng, Eric Smith, Ivan Evtimov, Song Jiang, et al.
Published
Oct 2025
Citations
10
Code
Not linked
32

arXiv.org

Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation

The paper introduces specification alignment, a challenge for LLMs to follow dynamic, scenario-specific behavioral and safety specifications. To evaluate this, the authors present SPECBENCH, a benchmark covering 5 scenarios, 103 specifications, and 1,500 prompts. Experiments on 33 models reveal significant alignment gaps and a safety-behavior trade-off.…

Haoran Zhang, Yafu Li, Xuyang Hu, Dongrui Liu, et al.
Published
Sep 2025
Citations
3
Code
24 stars
33

OpenAI

Why Language Models Hallucinate

The paper argues that language model hallucinations arise from statistical pressures during pretraining and persist due to misaligned evaluation metrics. The authors formalize hallucinations as errors in binary classification, showing that even with error-free training data, the cross-entropy objective leads to errors. They introduce the Is-It-Valid (IIV)…

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang
Published
Sep 2025
Citations
282
Code
Not linked
34

AAAI Conference on Artificial Intelligence

Beyond Transcription: Mechanistic Interpretability in ASR

This paper adapts interpretability methods from LLMs—logit lens, linear probing, and activation patching—to analyze the internal mechanisms of ASR models, specifically Whisper-large-v3 and Qwen2-Audio. The authors find that acoustic and semantic attributes (e.g., speaker gender, noise, accent) are linearly decodable from encoder layers, with peak…

Neta Glazer, Yael Segal-Feldman, Hilit Segev, Aviv Shamsian, et al.
Published
Aug 2025
Citations
11
Code
Not linked
35

arXiv.org

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models

The paper introduces FinCDM, the first cognitive diagnosis evaluation framework for financial large language models (LLMs), moving beyond aggregate scores to assess knowledge-skill level proficiency. It constructs CPA-KQA, a dataset of 210 expert-annotated questions derived from the CPA exam, covering 70 financial concepts, with high inter-annotator…

Ziyan Kuang, Feiyu Zhu, Maowei Jiang, Yanzhao Lai, et al.
Published
Aug 2025
Citations
4
Code
3 stars