AgentDish directory

AI Safety

Accepted listings with this tag.

Listing Category Score Trend Checked
#1202 ↓ -2
Evidence-to-Skill

A GitHub repository that turns untrusted source material into compact AI skills using evidence gates, validation steps, and audit output.

Developer Tools / AI Safety / Agent Skills 82 ↓ -2 30 days ago Details
#1493 ↑ +2
Open-AGS Attestor

A repository for a zero-trust execution boundary that sits between AI agents and real system actions. It evaluates policy, approval, scope, freshness, replay, and evidence before allowing operations to proceed.

Developer Tools / AI Safety / Governance 79 ↑ +2 51 days ago Details

A safety-evaluation harness for AI agents operating physical equipment. It validates proposed robot intent against declared hardware and deployment limits, records machine-readable refusal reasons, and includes a sample corpus plus a runnable evaluation script.

Developer Tools / AI Safety 78 ↑ +6 3 days ago Details

An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX.

Research / AI Safety 77 → 0 70 days ago Details

arXiv paper on how ideas and goals can spread through multi-agent LLM systems, including experiments on propagation, susceptibility factors, and mitigation.

Research / AI Safety 72 ↑ +1 11 days ago Details