AgentDish directory

mechanistic interpretability

Accepted listings with this tag.

Listing Category Score Trend Checked
#1545 ↑ +6
HyperSAE

HyperSAE is a Python package for mechanistic interpretability that trains hyperbolic sparse autoencoders on LLM activations. The page includes installation steps, a quickstart example, benchmark tables, and a short software architecture overview.

Developer Tools / AI/ML Frameworks 78 ↑ +6 20 days ago Details

An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX.

Research / AI Safety 77 → 0 70 days ago Details