CompletionKit by completionkit.com

MCP server · Agent observability & evals · indexed, not reviewed

Hostedvendor's own

Not reviewed

No score, grade or rank. This listing is facts from the official MCP registry and our own checks, and it stays out of the rankings until the panel reviews it.

How the index works

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.

What the official MCP registry says

Facts

MCP registry
com.completionkit/evals · 1.0.0
Endpoint
https://completionkit.com/mcp
GitHub stars
3
Registry entry
updated 18 Jul 2026

From the official MCP registry, the package registries and our own checks. JSON · Markdown

Why it's listed

  • It's published in the registry under completionkit.com, a namespace the registry only gives to whoever proves they control that domain.

Being indexed says nothing about quality, and nobody can pay for it. Is this yours? Ask for a review.

Tools it lists 54 · about 5,722 tokens of context · checked 20 hours ago

ToolWhat it doesHint
prompts_listList all prompts
prompts_getGet a prompt by ID
prompts_createCreate a prompt
prompts_updateUpdate a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don't go live without…
prompts_deleteDelete a prompt
prompts_publishPublish a prompt version, making it the current version
prompts_suggest_improvementSuggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (preserving…
runs_listList all runs
runs_getGet a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the metric dragging a…
runs_createCreate a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones.
runs_updateUpdate a run
runs_deleteDelete a run
runs_generateStart a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it.
runs_regradeRe-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run.
runs_rerunCreate and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions.
runs_retry_failuresRe-run only the failed responses of a run, optionally limited to specific response ids via "only".
responses_listList responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, "min_score"/"max_score" to isolate low…
responses_getGet a specific response
datasets_listList all datasets
datasets_getGet a dataset by ID
datasets_createCreate a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and to checks that compare against the…
datasets_updateUpdate a dataset
datasets_deleteDelete a dataset
datasets_create_from_urlCreate a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass through the tool-call…
metrics_listList all metrics
metrics_getGet a metric by ID
metrics_createCreate a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected…
metrics_updateUpdate a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max…
metrics_deleteDelete a metric
metrics_suggest_variantsAsk the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". Returns the persisted drafts.…
metric_groups_listList all metric groups
metric_groups_getGet a metric group by ID
metric_groups_createCreate a metric group
metric_groups_updateUpdate a metric group
metric_groups_deleteDelete a metric group
metric_versions_listList every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps.
metric_versions_publishPublish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, demotes peers, and writes the version's…
metric_versions_dismissDestroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as current instead.
provider_credentials_listList all provider credentials (API keys are not exposed)
provider_credentials_getGet a provider credential by ID (API key is not exposed)
provider_credentials_createCreate a provider credential
provider_credentials_updateUpdate a provider credential
provider_credentials_deleteDelete a provider credential
tags_listList all tags
tags_getGet a tag by ID
tags_createCreate a tag. Color is auto-assigned.
tags_updateRename a tag.
tags_deleteDelete a tag. Removes the tag from every linked metric, prompt, run, and dataset.
agreements_listList agreements. Filter by run_id, response_id, metric_id, or created_by.
agreements_createUpsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'.
judges_replayCreate a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call runs_generate to actually re-judge the…
judges_compareCompare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unavailable for check metrics.
promptfoo_importImport a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is-json become deterministic check…
usage_getGet this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to pre-check quota before starting runs.…

What https://completionkit.com/mcp answered to tools/list, asked without credentials. answered without the initialize handshake. The token figure is the size of the list as sent, divided by four; a model sees about that much before it calls anything. Full definitions, input schemas included, are in the listing's JSON under mcpTools.

How its tools read to an agent 0 errors · 124 warnings · 1 note

  • warnTC06datasets_deletethe description is "Delete a dataset"
  • warnTC06datasets_getthe description is "Get a dataset by ID"
  • warnTC06datasets_listthe description is "List all datasets"
  • warnTC06datasets_updatethe description is "Update a dataset"
  • warnTC06metric_groups_createthe description is "Create a metric group"
  • warnTC06metric_groups_deletethe description is "Delete a metric group"
  • warnTC06metric_groups_getthe description is "Get a metric group by ID"
  • warnTC06metric_groups_listthe description is "List all metric groups"
  • warnTC06metric_groups_updatethe description is "Update a metric group"
  • warnTC06metrics_deletethe description is "Delete a metric"
  • warnTC06metrics_getthe description is "Get a metric by ID"
  • warnTC06metrics_listthe description is "List all metrics"
  • warnTC06prompts_createthe description is "Create a prompt"
  • warnTC06prompts_deletethe description is "Delete a prompt"
  • warnTC06prompts_getthe description is "Get a prompt by ID"
  • warnTC06prompts_listthe description is "List all prompts"
  • warnTC06responses_getthe description is "Get a specific response"
  • warnTC06runs_deletethe description is "Delete a run"
  • warnTC06runs_listthe description is "List all runs"
  • warnTC06runs_updatethe description is "Update a run"
  • warnTC06tags_getthe description is "Get a tag by ID"
  • warnTC06tags_listthe description is "List all tags"
  • warnTC06tags_updatethe description is "Rename a tag."
  • warnTC11agreements_createnone of its 7 parameters has a description

The first 24 of 125; every finding is in the listing's JSON under mcpTools.check.

The checks from /check and anchor check, run each day on the list above: about 5,722 tokens of definitions. Not part of the score yet. Check your own server.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.