Projects
Here are some ongoing research directions of the group. Openings for curious students at the undergraduate and graduate levels!
How does information flow through transformers?
Transformers aggregate information from many pieces – language fragments, image patches, temporal intervals, etc. – into a global representation of the whole. Most commonly, they operate on point-based representations and no limitation on the amount of information passing through each latent space. What if we want to know what information is processed by each component at each location?
In this project, we convert specific junctures in the transformer into probabilistic representation spaces and restrict the transmission of information. One point-based processing path from input to output then becomes an information flow, with volume and mistakes and rates contributed by each chokepoint. We’re particularly interested in light-touch modifications to pre-trained models – see our recent stochastic normalization work below – and then in applying methods from mechanistic interpretability for a new perspective on previously identified circuitry.
Research highlights:
What’s the structure of information in data?

Data comes packaged in groups of features, generally with rich multivariate structure concerning their shared variation and the information relevant to downstream tasks. What if, while optimizing a model to process data, we also optimized the selection of information from the features? With a finite budget, we’d find an intuitive notion of the most important information for a task.
We could also probe the specific way that information is contained within the features – whether they are columns in a tabular dataset or pixels in an image. In the gif to the right, our recent Reveal-IG method turns sensitivity to information revelation into a heatmap of evidence for a ResNet-50 to classify each frame as a goldfish; the method is a hybrid between Shapley additive explanations (SHAP) and Integrated Gradients (IG), two of the most popular methods in Explainable AI.
Research highlights:
- Attribution via Distributional Paths for Information Revelation, arXiv 2026.
- Surveying the Space of Descriptions of a Composite System with Machine Learning, Physical Review Letters 2025.
- Information decomposition in complex systems via machine learning, PNAS 2024.
- Where is the information in data?, IEEE VIS 2024 workshop, Visualization for AI Explainability (visXAI).
- Interpretability with full complexity by constraining feature information, ICLR 2023.
How is information organized in probabilistic representation spaces?

We clearly like probabilistic representation spaces in this group. They can make information transmission finite and measurable, and they provide a notion of representational similarity that’s grounded in downstream computation (through the data processing inequality). In this project, our questions are more fundamental about the organization of information in such spaces. One curiosity arose from our study about compressing chaotic dynamical systems’ state spaces: although the neural network was optimized with a soft latent space, the eventual solution often hardened into a compression scheme equivalent to a discrete codebook.
Hard compression schemes are generally more practical and far more interpretable, while the space of soft schemes is far larger and amenable to optimization. Can we get the best of both worlds, through a better understanding of the space of compression schemes?
Research highlights:
- Comparing the information content of probabilistic representation spaces, TMLR 2025.
- Machine-Learning Optimized Measurements of Chaotic Dynamical Systems via the Information Bottleneck, Physical Review Letters 2024.