Do you sell AI governance or compliance software?Founding Vendor Placement — $499 for 12 months →
EU AI Act · Article 4 active · Article 50 from 2 Aug 2026 · high-risk rules transition through 2027–2028
Supporting capability — scope dependentSupporting governance capability

Testing, Evaluation & Observability

Testing, evaluation and monitoring capabilities can support several AI Act obligations, but a tool capability should not be treated as proof of legal compliance.

WHAT TO REVIEW

Practical checkpoints

  • Prompt injection & jailbreak red-teaming
  • Token tracing and latency monitoring
  • Model drift and bias alerting

Navigator mappings are a discovery aid. They do not establish that an organisation or system is legally in scope, compliant or certified.

HOW TO USE THIS PAGE

Start with the evidence

These software matches come from the requirement tags recorded in the Navigator catalogue. Review each vendor profile and its source evidence before making a procurement or compliance decision.

Run assessment
MAPPED SOFTWARE

Tools recorded for this requirement

50 source-linked profiles currently mapped to this requirement.

Giskard

AI Evaluation & Testing

Open-source testing platform backed by European Commission grants for LLM red-teaming and automated vulnerability scanning.

Paris, France€0 / €490/moNavigator highlight

LatticeFlow AI

Technical Documentation

Co-creator of the COMPL-AI framework evaluating foundation models directly against EU statutory mandates.

Zurich, Switzerland$1,200/moNavigator highlight

Promptfoo

AI Security & Guardrails

Open-source developer CLI with automated red-teaming, prompt injection benchmarks, and CI/CD security assertions.

San Francisco, USA$0 / $299/mo

Langfuse

AI Evaluation & Testing

Berlin-based open-source LLM observability platform providing continuous execution logging and trace audits in Frankfurt.

Berlin, Germany$0 / $59/mo

Lakera

AI Security & Guardrails

Swiss AI security company providing real-time inference firewalls against prompt injection, jailbreaks, and PII leaks.

Zurich, Switzerland$800/mo

LangSmith

AI Evaluation & Testing

Developer-centric LLM observability platform providing automated tracing, regression evaluation, and runtime prompt logging.

San Francisco, USA$0 / $39/mo

Fiddler AI

AI Evaluation & Testing

Enterprise observability platform providing explainable AI (XAI), fairness metrics, and production drift telemetry.

Palo Alto, USA$1,200/mo

Arize AI

AI Evaluation & Testing

AI evaluation and observability platform tracking embedding drift, real-time evals, and prompt telemetry.

San Francisco, USA$0 / $500/mo

WhyLabs

AI Evaluation & Testing

Privacy-preserving AI observability platform computing statistical summaries on-device without data leakage.

Seattle, USA$0 / $450/mo

TruEra

AI Evaluation & Testing

AI quality management suite evaluating explainability, relevance, and hallucination scores across LLMs and predictive ML.

Redwood City, USA$800/mo

Encord

Technical Documentation

Data development platform providing automated annotation quality control, data lineage, and bias curation.

London, UK$600/mo

Kolena

AI Evaluation & Testing

Scenario-driven model evaluation platform verifying model behavior across granular sub-cohorts and edge cases.

San Francisco, USA$1,000/mo

Armilla AI

AI Governance & Inventory

AI assurance engine providing automated risk verification and performance warranties for enterprise AI models.

Toronto, Canada$1,600/mo

Arthur AI

AI Evaluation & Testing

Model monitoring and governance engine offering real-time prompt risk scoring, hallucination tracking, and bias evaluation.

New York, USA$1,250/mo

Deepchecks

AI Evaluation & Testing

Continuous AI validation platform providing automated test suites across data integrity, model behavior, and LLM evaluation.

Tel Aviv, Israel$0 / custom pricing

Patronus AI

AI Evaluation & Testing

Automated LLM evaluation platform specializing in hallucination scoring, copyright exposure detection, and enterprise safety benchmarking.

New York, USA$900/mo

Galileo

AI Evaluation & Testing

Evaluation and guardrails platform providing automated chain-of-thought metrics, prompt debugging, and production safety scoring.

San Francisco, USA$500/mo

Braintrust

AI Evaluation & Testing

Enterprise-grade evaluation engine providing fast automated testing, scoring loops, and production AI telemetry.

San Francisco, USA$0 / $249/mo

Securiti

AI Governance & Inventory

DataAI security and governance platform that discovers and catalogs AI models and agents, classifies AI risk, maps data-to-AI relationships, and assesses AI systems against the EU AI Act and other regulations.

San Jose, USAContact sales

Qualdo

AI Evaluation & Testing

Continuous ML and data quality monitoring platform providing automated anomaly detection, performance tracking, and confidence scoring.

San Jose, USA$450/mo

Aimon AI

AI Evaluation & Testing

Continuous observability and safety API evaluating LLM hallucination, instruction drift, and factual alignment.

San Francisco, USA$0 / $199/mo

DeepKeep

AI Security & Guardrails

AI security and robustness platform running automated white-box and black-box penetration tests across vision and language models.

Tel Aviv, Israel$1,500/mo

Enkrypt AI

AI Security & Guardrails

AI security and compliance platform providing red teaming, runtime guardrails, policy enforcement, monitoring and audit-ready evidence.

Brighton, USA$0 / $149/mo

Unify AI

AI Evaluation & Testing

Dynamic LLM routing infrastructure ensuring high availability, latency optimization, and automated failover compliance.

London, UK$0 / Usage

Trustwise

AI Security & Guardrails

Safety framework and proxy for autonomous AI agents, preventing unauthorized tool calls, hallucinations, and policy breaches.

Austin, USA$950/mo

Relari AI

AI Evaluation & Testing

Synthetic dataset generation and RAG evaluation platform stress-testing generative pipelines on enterprise edge cases.

San Francisco, USA$0 / $400/mo

ModelScan

AI Security & Guardrails

Open-source serialization vulnerability scanner inspecting PyTorch, Keras, and Pickle model weights for code execution exploits.

Seattle, USAOpen Source / $0

ClearML

Technical Documentation

Open-source MLOps platform recording full environment configurations, dataset versions, and training telemetry for model reproducibility.

Tel Aviv, Israel$0 / $490/mo

HiddenLayer

AI Security & Guardrails

AI security platform covering AI asset discovery, model scanning, red teaming and runtime protection across the AI lifecycle.

USAContact sales

EvalML (Alteryx)

AI Evaluation & Testing

Open-source AutoML library that builds, optimizes, and evaluates machine learning pipelines using domain-specific objective functions.

Irvine, USAOpen Source / $0

Fairlearn

AI Evaluation & Testing

Open-source Python framework providing fairness metrics, disparate impact mitigation algorithms, and comparative fairness dashboards.

Redmond, USAOpen Source / $0

Context.ai

AI Evaluation & Testing

Product analytics and user interaction tracking platform for LLM applications, monitoring conversation drop-offs and sentiment drift.

London, UK$0 / $299/mo