Submission history

From: Matteo He

  1. Tue, 7 Jul 2026 11:26:16 UTC 68cbc05
  2. Thu, 3 Sep 2026 00:43:38 UTC fac8356
  3. Wed, 9 Sep 2026 22:33:45 UTC 0b2470b
  4. Thu, 10 Sep 2026 10:02:07 UTC fa69648
  5. Fri, 25 Sep 2026 13:44:49 UTC fc31c66
  6. Fri, 25 Sep 2026 17:08:24 UTC bba1628
  7. Fri, 25 Sep 2026 17:52:57 UTC e7dc632
  8. Fri, 25 Sep 2026 17:58:34 UTC e2dc873
  9. Fri, 25 Sep 2026 17:59:00 UTC 7995b99
  10. Fri, 25 Sep 2026 17:59:00 UTC 0aa6765
  11. Fri, 25 Sep 2026 17:59:00 UTC 98535c7
  12. Fri, 25 Sep 2026 19:25:47 UTC f60c7d7
  13. Fri, 25 Sep 2026 19:25:47 UTC d4e5043
  14. Fri, 25 Sep 2026 19:38:50 UTC 1d7d2b8
  15. Fri, 25 Sep 2026 19:38:50 UTC 1878266
  16. Fri, 25 Sep 2026 19:38:50 UTC bbf45fd
  17. Fri, 25 Sep 2026 19:38:50 UTC 23b4401
  18. Fri, 25 Sep 2026 20:56:46 UTC e0dbe2f
  19. Sat, 26 Sep 2026 16:53:37 UTC 45bbcea
  20. Sat, 26 Sep 2026 17:54:17 UTC c09818c
  21. Sat, 26 Sep 2026 18:00:09 UTC 8f5ec09
  22. Sat, 26 Sep 2026 18:21:57 UTC 6883bb7
  23. Sat, 26 Sep 2026 18:31:38 UTC d076844
  24. Sat, 26 Sep 2026 23:20:55 UTC 836b809
  25. Sat, 26 Sep 2026 23:37:24 UTC 5d15ca1
  26. Sun, 27 Sep 2026 13:56:21 UTC edd5a00

Matteo He

∗

I’m a researcher interested in understanding AI systems and making their behavior easier to evaluate and oversee.

Abstract

I completed an MPhil in Advanced Computer Science at Cambridge with Distinction. My dissertation studied how language models learn to turn internal information into predictions. The resulting first-author paper, Learning to Read Out, was accepted to NeurIPS 2026.

I’m seeking AI safety research fellowships and research scientist or research engineer roles. Get in touch →

Correspondence to matteohe.research@gmail.com.

Recent updates

  1. Learning to Read Out accepted at NeurIPS 2026.

  2. Sparse Readout Prism preprint released on arXiv.

  3. Completed the MPhil in Advanced Computer Science at Cambridge with Distinction.

Selected research

Learning to Read Out

Unembedding Dynamics in Language Model Pretraining

Matteo He, William F. Shen, Alex Iacob, Andrej Jovanovic, Xinchi Qiu, Nicholas D. Lane

First authorNeurIPS 2026 · Accepted

When does a language model acquire information, and when can it use that information to predict the next word?

A model’s output layer turns its internal information into scores for possible next words. I studied how that layer changes during training, using probes to check what information was already present and trajectory crosscoders to track how the output weights developed.

In subject–verb agreement tests, grammatical information was detectable in hidden states hundreds of training steps before the model used it reliably in its predictions. The gap remained on harder sentences at the end of training. Predictions alone can therefore give a misleading date for when this information became available.

Explore how readout features develop during training →
Measured feature trajectories in Pythia-1B, normalized to each feature’s peak, show different patterns of change during training. Open the explorer to inspect all four models.
One model from the study: 1,200 sampled features, each scaled to its own peak. Explore all four models →
Method & scope

I tracked snapshots of the output weights during pretraining and fitted trajectory crosscoders to study how they changed. I tested the resulting features with readout swaps, linear probes, and ablations. On harder examples involving relative clauses, the gap between linear decodability and the model’s own readout persisted at convergence.

Subject–verb agreement results in Pythia-6.9B. Left: a probe of hidden states reaches high accuracy before the model’s own readout. Right: at convergence, probe accuracy remains at ceiling while the model’s readout falls behind on harder grammatical dependencies, with the largest gap at 25 percentage points.
Left: grammatical information becomes detectable before the model’s own readout uses it reliably. Right: the gap persists on harder examples, even at the end of training. SVA means subject–verb agreement.

Using fixed hidden states from Pythia-1B at step 1000, removing eight localized features from the reconstructed final readout reduced agreement accuracy on held out examples from 90.0% to 48.7%. Matched controls stayed at 90.0%. In control runs with 31 million parameters and a single random seed, each 4× increase in the readout learning rate halved the training step at which reorganization occurred. Final validation losses agreed within 0.06 nats.

The public release includes reusable analysis tools, examples that run on a CPU, guided notebooks, and a manifest linking experiments to figures. The paper was accepted to NeurIPS 2026.

Sparse Readout Prism

Explaining Logit-Lens Scores in Features Instead of Tokens

Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane

First authorPreprint · Under review

What do token scores leave out about a model’s internal representations?

A lens translates model internals into token scores, but its fitting data can shape what it reports. In a controlled test, lenses fitted on English and Chinese reported different languages for the same hidden states. I developed Sparse Readout Prism to explain these scores through shared features of the output weights, with an explicit account of what remains unexplained.

Explore what token scores leave out →
Sparse readout features Signed score accounting 𝑞𝛼=∑𝑖𝛽𝑖(𝛼)𝑑𝑖+𝑟𝛼 𝑠𝛼=∑𝑖𝑐𝑖+ℎ̃⊤ℓ𝑟𝛼 𝑞𝛼 𝛽𝑖(𝛼)𝑑𝑖 𝑟𝛼 B 0 A ℎ̃⊤ℓ𝑟𝛼
How individual features contribute to a readout score.
Method & scope

The method fits sparse autoencoders directly to the model’s output weights and decomposes selected logits and margins into signed feature contributions. Across six readouts without softcaps, it improved reconstruction coverage for logit differences by 8.9 to 17.3 percentage points over the strongest of six geometric baselines. A reconstruction counts as covered when it has the correct sign and less than 50% relative error. I use reconstruction and replacement tests, alongside checks for sign errors, residuals, and confounds, to establish the explanation’s limits.

Two paper examples show signed feature differences between insect and software senses of bug, and structure and network senses of bridge.
Selected feature contrasts from the paper: “bug” and “bridge”.

Education & distinctions

University of Cambridge

MPhil, Advanced Computer Science

Distinction · AI Alignment Fellowship

University of St Andrews

BSc (Hons), Computer Science & Mathematics

First Class · 18/20 average · Dean’s List in all three years

Top Student Medal for the highest academic achievement in the Direct Entry Computer Science cohort.

GRE: 340/340Quantitative Reasoning: 170/170 · Verbal Reasoning: 170/170

Experience & projects

Amazon Alexa AI

Software Development Engineer Intern

I built ML data and experiment infrastructure used by 20+ Applied Scientists. The work supported faster dataset iteration and AWS pipelines processing millions of utterances daily.

My focus was infrastructure for model evaluation within the Alexa NLU team.

vigil-gpu

Open source · Python

A terminal monitor for rented cloud GPUs. Tracks training logs and metrics across instances, with reconnecting SSH streams and alerts for stalled runs, NaNs, and loss plateaus.

LLMs in the browser

Independent project · Archived

I built and operated a platform for learning Japanese that served approximately 3,000 monthly active users. WebGPU/WebLLM ran inference on users’ devices, eliminating server inference costs.

Engineering scope

I built and maintained the full system, including inference, frontend and backend, authentication, subscriptions, deployment, and monitoring. The platform is now archived and maintenance is paused.

Apps

3D Game of Life

Web app · three.js, WebGPU

Conway’s Game of Life on a 3D grid of up to 100³ cells. A cell has 26 neighbours rather than 8, so the app is built around finding rules that work: set the birth and survival counts, weight neighbours by face, edge and corner, or let the rule finder try hundreds of random rules and keep the ones that stay alive. It runs on the GPU through WebGPU, and any world can be shared as a short link or viewed in VR.

A cube of white voxels, grown by 3D Game of Life rules, inside a wireframe box on a deep blue background.
Drag to orbit, scroll to zoom, space to run.

好牌 · Good Hand

Web app · Python server, browser client

Dou Dizhu (斗地主, “fight the landlord”) with friends and bots, in English, Chinese and Italian. It has the classic three-player game and a four-player, two-deck variant with a secret partner, plus a guided first round and ranked play. The strong bots run DouZero, a self-play RL policy (ICML 2021), which I ported to numpy and taught to bid from its own value of the opening hand.

A Dou Dizhu table: the bot landlord has played a straight from ten to ace, marked with the characters 顺子, above the player’s seventeen-card hand.
The bot landlord leads a straight, 10 to A.

Small apps

Matrix Digital Rain

Interactive · One HTML file, no dependencies

The falling code from The Matrix as a corridor you fly through, steering with the mouse or a finger. The space bar warps it into a tunnel, and the film’s flat columns are one key away. Each view grows from a seed, such as a name, kept in the page address so it can be shared as a link.

Green glyphs falling down the walls of a corridor that recedes into darkness.
Steer with the mouse or a finger; space to warp.

Matrix Cave

Interactive · One HTML file, no dependencies

A generated cave lit only by the falling code: each string is a drop of water running over the rock, so the walls, pools and lava show only where it traces them. You walk through at a person’s pace, and E sends out a pulse that lights the passages. Every sound is synthesised in the page, with drips delayed and muffled by distance and rock. Caves are seeded and shareable as links.

Green glyphs running down the walls and floor of a cave passage that recedes into darkness.
Click to look, WASD to walk, E to ping.

Contact

For enquiries about research roles, fellowships, or collaborations.

Contact me on LinkedIn

matteohe.research@gmail.com

Research figure