Event date · · IceCube

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

FACT STATEMENT

A study applies sparse-autoencoder-based mechanistic interpretability to a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction. It identifies a validated atlas of physical concepts using held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show the direction head barely uses this atlas. An uncertainty head trained on the same event-level representation to predict angular reconstruction error depends causally on quality and brightness features from the atlas. At 20% selection efficiency, the interpretable estimator improves median angular resolution from 20.2° to 3.2°.

What happened

Researchers present the first application of sparse-autoencoder-based mechanistic interpretability to particle physics. They study a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction. They identify a validated atlas of physical concepts in the model representation using a strict validation protocol. Causal interventions show the direction head barely draws on this atlas. Motivated by this underused information, they train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At 20% selection efficiency, this interpretable estimator improves the median angular resolution from 20.2° to 3.2°. The results suggest mechanistic interpretability can reveal learned latent physics and help design downstream tasks that exploit it.

Technical significance

Sparse autoencoders are used to extract interpretable latents from a neutrino foundation model. The validation protocol includes held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions reveal that the direction head does not rely on the identified atlas, while an uncertainty head trained on the same representation does depend on quality and brightness features. The uncertainty head improves median angular resolution from 20.2° to 3.2° at 20% selection efficiency.

Industry impact

This work demonstrates that mechanistic interpretability can be applied beyond language models to scientific foundation models, potentially enabling better understanding and control of learned representations in physics. The approach may be transferable to other domains where foundation models are used for scientific data analysis.

Decision value

The method could lead to more accurate and interpretable models for particle physics experiments, potentially improving data analysis efficiency and enabling new scientific discoveries. The uncertainty estimation approach may be valuable for selecting high-quality events in large-scale experiments.

What to watch

Future work may explore applying similar interpretability techniques to other scientific foundation models, and further investigate how identified latents can be leveraged to improve downstream task performance. The underuse of the atlas by the direction head suggests potential for redesigning model heads to better exploit learned features.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.