CompletionKit by completionkit.com
MCP server · Agent observability & evals · indexed, not reviewed
Hostedvendor's own
Not reviewed
No score, grade or rank. This listing is facts from the official MCP registry and our own checks, and it stays out of the rankings until the panel reviews it.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Facts
- MCP registry
com.completionkit/evals· 1.0.0- Endpoint
https://completionkit.com/mcp- Website
- completionkit.com
- GitHub stars
- 3
- Registry entry
- updated 18 Jul 2026
From the official MCP registry, the package registries and our own checks. JSON · Markdown
Why it's listed
- It's published in the registry under completionkit.com, a namespace the registry only gives to whoever proves they control that domain.
Being indexed says nothing about quality, and nobody can pay for it. Is this yours? Ask for a review.
Tools it lists 54 · about 5,722 tokens of context · checked 20 hours ago
| Tool | What it does | Hint |
|---|---|---|
prompts_list | List all prompts | |
prompts_get | Get a prompt by ID | |
prompts_create | Create a prompt | |
prompts_update | Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don't go live without… | |
prompts_delete | Delete a prompt | |
prompts_publish | Publish a prompt version, making it the current version | |
prompts_suggest_improvement | Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (preserving… | |
runs_list | List all runs | |
runs_get | Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the metric dragging a… | |
runs_create | Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones. | |
runs_update | Update a run | |
runs_delete | Delete a run | |
runs_generate | Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it. | |
runs_regrade | Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run. | |
runs_rerun | Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions. | |
runs_retry_failures | Re-run only the failed responses of a run, optionally limited to specific response ids via "only". | |
responses_list | List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, "min_score"/"max_score" to isolate low… | |
responses_get | Get a specific response | |
datasets_list | List all datasets | |
datasets_get | Get a dataset by ID | |
datasets_create | Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and to checks that compare against the… | |
datasets_update | Update a dataset | |
datasets_delete | Delete a dataset | |
datasets_create_from_url | Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass through the tool-call… | |
metrics_list | List all metrics | |
metrics_get | Get a metric by ID | |
metrics_create | Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected… | |
metrics_update | Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max… | |
metrics_delete | Delete a metric | |
metrics_suggest_variants | Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". Returns the persisted drafts.… | |
metric_groups_list | List all metric groups | |
metric_groups_get | Get a metric group by ID | |
metric_groups_create | Create a metric group | |
metric_groups_update | Update a metric group | |
metric_groups_delete | Delete a metric group | |
metric_versions_list | List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps. | |
metric_versions_publish | Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, demotes peers, and writes the version's… | |
metric_versions_dismiss | Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as current instead. | |
provider_credentials_list | List all provider credentials (API keys are not exposed) | |
provider_credentials_get | Get a provider credential by ID (API key is not exposed) | |
provider_credentials_create | Create a provider credential | |
provider_credentials_update | Update a provider credential | |
provider_credentials_delete | Delete a provider credential | |
tags_list | List all tags | |
tags_get | Get a tag by ID | |
tags_create | Create a tag. Color is auto-assigned. | |
tags_update | Rename a tag. | |
tags_delete | Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset. | |
agreements_list | List agreements. Filter by run_id, response_id, metric_id, or created_by. | |
agreements_create | Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'. | |
judges_replay | Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call runs_generate to actually re-judge the… | |
judges_compare | Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unavailable for check metrics. | |
promptfoo_import | Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is-json become deterministic check… | |
usage_get | Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to pre-check quota before starting runs.… |
What https://completionkit.com/mcp answered to tools/list, asked without credentials. answered without the initialize handshake. The token figure is the size of the list as sent, divided by four; a model sees about that much before it calls anything. Full definitions, input schemas included, are in the listing's JSON under mcpTools.
How its tools read to an agent 0 errors · 124 warnings · 1 note
- warnTC06datasets_deletethe description is "Delete a dataset"
- warnTC06datasets_getthe description is "Get a dataset by ID"
- warnTC06datasets_listthe description is "List all datasets"
- warnTC06datasets_updatethe description is "Update a dataset"
- warnTC06metric_groups_createthe description is "Create a metric group"
- warnTC06metric_groups_deletethe description is "Delete a metric group"
- warnTC06metric_groups_getthe description is "Get a metric group by ID"
- warnTC06metric_groups_listthe description is "List all metric groups"
- warnTC06metric_groups_updatethe description is "Update a metric group"
- warnTC06metrics_deletethe description is "Delete a metric"
- warnTC06metrics_getthe description is "Get a metric by ID"
- warnTC06metrics_listthe description is "List all metrics"
- warnTC06prompts_createthe description is "Create a prompt"
- warnTC06prompts_deletethe description is "Delete a prompt"
- warnTC06prompts_getthe description is "Get a prompt by ID"
- warnTC06prompts_listthe description is "List all prompts"
- warnTC06responses_getthe description is "Get a specific response"
- warnTC06runs_deletethe description is "Delete a run"
- warnTC06runs_listthe description is "List all runs"
- warnTC06runs_updatethe description is "Update a run"
- warnTC06tags_getthe description is "Get a tag by ID"
- warnTC06tags_listthe description is "List all tags"
- warnTC06tags_updatethe description is "Rename a tag."
- warnTC11agreements_createnone of its 7 parameters has a description
The first 24 of 125; every finding is in the listing's JSON under mcpTools.check.
The checks from /check and anchor check, run each day on the list above: about 5,722 tokens of definitions. Not part of the score yet. Check your own server.