WebMCP Challenge 2026

MetaFence

Agents can triage metadata without ingesting the raw strings that may try to manipulate them. Humans keep the final review and export controls.

Checking WebMCP…

Evaluator quick path · about 60 seconds

Test the WebMCP boundary, not just the interface

  1. Ask the agent to summarize_metadata_risk. It receives counts and reason codes—not raw descriptions.
  2. Ask it to stage_metadata_review with {"risk":"quarantine","maxRecords":3}. Three bounded records appear visibly below.
  3. Ask it to request_human_metadata_review for instruction-001. The page opens raw text for the human, while the tool result still withholds it.
  4. Ask it to stage_safe_csv_export with {"risk":"quarantine"}. The agent can stage four safe structural rows; only the human can download.

What this proves: non-trivial WebMCP leverage · complete human/agent execution · a concrete prompt-injection boundary · a least-authority pattern that differs from generic click automation.

A safer human + agent boundary

Untrusted metadata→Deterministic scanner→Safe agent tools→Visible human decision

Tool responses include IDs, lengths, risk labels, and reasons—never raw descriptions. An agent can open one record in the page for a human, but cannot approve, rewrite, publish, or silently download anything.

Risk summary

Human review queue

Eight synthetic records demonstrate clear, review, and prompt-injection quarantine paths.

IDPageRiskDescription lengthReasons

Bounded export

Stage a CSV containing only safe structural fields. Download still requires a human click.

Visible event log