AgentDish directory
mechanistic interpretability
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1545
↑ +6
HyperSAE
HyperSAE is a Python package for mechanistic interpretability that trains hyperbolic sparse autoencoders on LLM activations. The page includes installation steps, a quickstart example, benchmark tables, and a short software architecture overview. |
Developer Tools / AI/ML Frameworks | 78 | ↑ +6 | 20 days ago | Details |
|
#1652
→ 0
Prompt Injection as Role Confusion
An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX. |
Research / AI Safety | 77 | → 0 | 70 days ago | Details |