Open source

  • Assevra AI

    An MIT-licensed, open-source reliability scorecard for LLM agents. Assevra AI scores agent reliability against fixed thresholds so teams can measure — not guess — whether an agent is safe to ship.

    Founder and maintainer.

    • Scores agent reliability across four dimensions: grounding and faithfulness, safety and refusal, PII-leak detection, and task completion.
    • Evaluates against fixed thresholds with 95% Wilson confidence intervals for statistically honest pass/fail signals.
    • Ships as a CLI with Markdown, JSON, and HTML reports plus a CI gate for continuous reliability checks.
    • Written in Python and published as an installable package on PyPI.
    • Python
    • LLM evaluation
    • CLI
    • CI/CD

Published writing