AI Developer Tool / Benchmarking / evaluation

TaxCalcBench pull request: TY25 Kimi K3 support via OpenRouter

Pull request adding TY25 support for moonshotai/kimi-k3 in TaxCalcBench through OpenRouter. The snapshot shows benchmark wiring, PDF handling details, saved outputs and evaluation reports, and regenerated charts/results for the full TY25 set.

Clear22/30
Useful26/30
Specific12/20
Complete16/20
TaxCalcBench pull request: TY25 Kimi K3 support via OpenRouter screenshot

Why it was accepted

The page clearly describes an AI-related developer benchmark update with concrete implementation and validation evidence: OpenRouter integration, model-specific reasoning settings, PDF input handling, saved outputs, evaluation reports, and test/chart regeneration. It is specific enough to be useful as a public listing and is clearly tied to an AI benchmark project rather than generic discussion.

Weakness

This snapshot is a merged pull request, so it does not fully show the surrounding repository structure, how to run the benchmark end to end, or the broader scope of TaxCalcBench beyond TY25/Kimi K3 support.

Review status

40 days ago #1689 ↓ -1

Last evaluated 40 days ago. Current rank #1689. Down 1 spot in the rankings.

Score history

76

Related listings

prompt-scrub screenshot
90

AI Developer Tool / Prompt privacy / PII redaction

A local-first Node.js utility that redacts emails, phone numbers, paths, secrets, URLs, and other identifiers from LLM prompts, then rehydrates responses locally using session-based mappings.

ORBIT screenshot
#105 ORBIT
89

AI Developer Tool / AI Gateway / RAG Backend

Self-hosted, OpenAI-compatible AI gateway for private RAG, natural-language data access, and tool-calling agents. It connects files, databases, vector stores, APIs, and MCP tools through one backend with auth, observability, and governance features.

claude-code-eyes screenshot

AI Developer Tool / Vision / Hardware Verification

A Claude Code skill that lets the model inspect live camera frames so it can verify hardware displays, spot wiring issues, and check physical UI behavior on real devices.

state-harness screenshot
89

AI Developer Tool / LLM Agent Monitoring

Runtime safety net for LLM agents that detects runaway token growth, classifies failure patterns, and suggests fixes without extra LLM calls. The project includes a Rust core, Python SDK, examples, benchmarks, and benchmark claims across multiple models and runs.