AgentDish directory
AI Safety
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1202
↓ -2
Evidence-to-Skill
A GitHub repository that turns untrusted source material into compact AI skills using evidence gates, validation steps, and audit output. |
Developer Tools / AI Safety / Agent Skills | 82 | ↓ -2 | 30 days ago | Details |
|
#1493
↑ +2
Open-AGS Attestor
A repository for a zero-trust execution boundary that sits between AI agents and real system actions. It evaluates policy, approval, scope, freshness, replay, and evidence before allowing operations to proceed. |
Developer Tools / AI Safety / Governance | 79 | ↑ +2 | 51 days ago | Details |
|
#1534
↑ +6
URML physical-ai-safety-eval
A safety-evaluation harness for AI agents operating physical equipment. It validates proposed robot intent against declared hardware and deployment limits, records machine-readable refusal reasons, and includes a sample corpus plus a runnable evaluation script. |
Developer Tools / AI Safety | 78 | ↑ +6 | 3 days ago | Details |
|
#1652
→ 0
Prompt Injection as Role Confusion
An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX. |
Research / AI Safety | 77 | → 0 | 70 days ago | Details |
|
arXiv paper on how ideas and goals can spread through multi-agent LLM systems, including experiments on propagation, susceptibility factors, and mitigation. |
Research / AI Safety | 72 | ↑ +1 | 11 days ago | Details |