PlatPhorm Evals · Evidence-Backed QA

Prove your tools work.

Run evidence-backed evaluations across PlatPhormNews sites, APIs, MCP tools, browser journeys, schemas, workflows, and releases. Evals turns discovery, tests, traces, screenshots, sandbox runs, and model grades into scorecards and release decisions.

PlatPhorm Evals tests every site, API, MCP tool, workflow, UI, schema, and release path in the PlatPhormNews network, then gives humans and agents public-safe evidence, scorecards, and release gates.

Current readiness state

Degraded but usable

Database is configured but no registry services are persisted yet; Evals is using static fallback targets until registry sync writes records.

Latest real evidence

No persisted run available

This is an honest empty state. Use the launcher for public-safe dry-run evidence or run protected backend execution when provider evidence is required.

What Evals Does

Evaluation Mesh and Release-Control Mesh

Evals is the canonical quality, regression, evidence, scorecard, release-control, tool-validation, MCP/API evaluation, BrowserOps journey evaluation, Spec contract validation, Sandbox execution verification, AgentUI render validation, Claws orchestration evaluation, and LLM-as-judge platform for PlatPhormNews.

Step 1
Discover targets
Step 2
Generate suites
Step 3
Run checks
Step 4
Capture evidence
Step 5
Score results
Step 6
Gate releases
Step 7
Publish reports

What this service owns

scorecards
evidence grading
release gates
eval suites
findings
regression comparisons
public-safe readiness signals
confirmation URLs for Evals artifacts

What this service does not own

BrowserOps screenshots
Spec contract authoring
MCP registry mutation
Sandbox execution
Docs publishing
Sheets exports
Trace storage
AgentUI workflow orchestration

Who uses this?

humans reviewing releases
agents validating tool chains
developers testing APIs
operators checking service health
CI jobs blocking regressions
MCP clients validating tools
BrowserOps validating UI
Sandbox validating execution
Spec validating contracts

What gets evaluated?

APIs
MCP tools
OpenAPI schemas
AgentUI forms
BrowserOps journeys
Sandbox commands
Claws workflows
discovery files
policies
traces
RSS/sitemaps
route health
release readiness

Recent Runs

Latest persisted evaluation run results

View all runs
No eval runs are persisted yet. This is an honest empty state; no suite is being shown as passing.

Network Coverage

static fallback targets pending protected registry sync

View full registry
degraded
Avg Coverage
31
Services
996
Capabilities
static_fallback
Source
Network Integrations

Actionable integration status

Cards show persisted live status when synced. Static fallback targets are labeled pending sync and do not count as passing provider evidence.

View full matrix
pending_sync
Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

pending_sync
Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

pending_sync
Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Guided launch

Run a first public-safe discovery, OpenAPI, MCP, AgentUI, workflow, or CLI registry eval.

Evidence objects

Inspect public-safe artifacts, trace links, empty states, and redaction boundaries.

platphormctl

Use the CLI harness for repeatable discovery, MCP, policy, and dry-run validation.