# CompletionKit (slim) > CompletionKit, an MCP server by completionkit.com, listed from the official MCP registry. Indexed, not reviewed: facts and our own checks, no score or ranking. Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. - Full: https://www.anchorterminal.com/tools/completionkit-evals.md (~2,350 tokens) · this version ~2,280 tokens · JSON https://www.anchorterminal.com/tools/completionkit-evals.json · canonical https://www.anchorterminal.com/tools/completionkit-evals - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 # CompletionKit > Indexed, not reviewed: facts from the official MCP registry and our own checks. No score, grade or rank, and not in the rankings until the panel reviews it. How the index works: https://www.anchorterminal.com/indexed/ - Kind: MCP server, by completionkit.com (https://completionkit.com) - Category: Agent observability & evals (https://www.anchorterminal.com/categories/agent-observability.md) - Listed because: It's published in the registry under completionkit.com, a namespace the registry only gives to whoever proves they control that domain. - What the official MCP registry says: Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. ## Facts - MCP registry: `com.completionkit/evals` 1.0.0 - Endpoint: https://completionkit.com/mcp (streamable HTTP) - Source: https://github.com/homemade-software-inc/completion-kit - Website: https://completionkit.com - GitHub stars: 3 - Registry entry updated: 2026-07-18 ## Tools - Tools it lists (54, about 5,722 tokens of context, `tools/list` without credentials, checked 2026-10-04 22:22 UTC): - `prompts_list`: List all prompts - `prompts_get`: Get a prompt by ID - `prompts_create`: Create a prompt - `prompts_update`: Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with… - `prompts_delete`: Delete a prompt - `prompts_publish`: Publish a prompt version, making it the current version - `prompts_suggest_improvement`: Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns… - `runs_list`: List all runs - `runs_get`: Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and… - `runs_create`: Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones. - `runs_update`: Update a run - `runs_delete`: Delete a run - `runs_generate`: Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded… - `runs_regrade`: Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated… - `runs_rerun`: Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of… - `runs_retry_failures`: Re-run only the failed responses of a run, optionally limited to specific response ids via "only". - `responses_list`: List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields"… - `responses_get`: Get a specific response - `datasets_list`: List all datasets - `datasets_get`: Get a dataset by ID - `datasets_create`: Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth)… - `datasets_update`: Update a dataset - `datasets_delete`: Delete a dataset - `datasets_create_from_url`: Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV… - `metrics_list`: List all metrics - `metrics_get`: Get a metric by ID - `metrics_create`: Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value… - `metrics_update`: Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern… - `metrics_delete`: Delete a metric - `metrics_suggest_variants`: Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with… - `metric_groups_list`: List all metric groups - `metric_groups_get`: Get a metric group by ID - `metric_groups_create`: Create a metric group - `metric_groups_update`: Update a metric group - `metric_groups_delete`: Delete a metric group - `metric_versions_list`: List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps. - `metric_versions_publish`: Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'.… - `metric_versions_dismiss`: Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish… - `provider_credentials_list`: List all provider credentials (API keys are not exposed) - `provider_credentials_get`: Get a provider credential by ID (API key is not exposed) - `provider_credentials_create`: Create a provider credential - `provider_credentials_update`: Update a provider credential - `provider_credentials_delete`: Delete a provider credential - `tags_list`: List all tags - `tags_get`: Get a tag by ID - `tags_create`: Create a tag. Color is auto-assigned. - `tags_update`: Rename a tag. - `tags_delete`: Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset. - `agreements_list`: List agreements. Filter by run_id, response_id, metric_id, or created_by. - `agreements_create`: Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is… - `judges_replay`: Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only… - `judges_compare`: Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong… - `promptfoo_import`: Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics;… - `usage_get`: Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period… - How its tools read to an agent (0 errors, 124 warnings, 1 note, about 5,722 tokens; rules at https://www.anchorterminal.com/check.md; not part of the score): - warn TC06 datasets_delete: the description is "Delete a dataset" - warn TC06 datasets_get: the description is "Get a dataset by ID" - warn TC06 datasets_list: the description is "List all datasets" - warn TC06 datasets_update: the description is "Update a dataset" - warn TC06 metric_groups_create: the description is "Create a metric group" - warn TC06 metric_groups_delete: the description is "Delete a metric group" - warn TC06 metric_groups_get: the description is "Get a metric group by ID" - warn TC06 metric_groups_list: the description is "List all metric groups" - warn TC06 metric_groups_update: the description is "Update a metric group" - warn TC06 metrics_delete: the description is "Delete a metric" - warn TC06 metrics_get: the description is "Get a metric by ID" - warn TC06 metrics_list: the description is "List all metrics" - warn TC06 prompts_create: the description is "Create a prompt" - warn TC06 prompts_delete: the description is "Delete a prompt" - warn TC06 prompts_get: the description is "Get a prompt by ID" - warn TC06 prompts_list: the description is "List all prompts" - warn TC06 responses_get: the description is "Get a specific response" - warn TC06 runs_delete: the description is "Delete a run" - warn TC06 runs_list: the description is "List all runs" - warn TC06 runs_update: the description is "Update a run" - warn TC06 tags_get: the description is "Get a tag by ID" - warn TC06 tags_list: the description is "List all tags" - warn TC06 tags_update: the description is "Rename a tag." - warn TC11 agreements_create: none of its 7 parameters has a description - JSON: https://www.anchorterminal.com/api/v1/tools/completionkit-evals.json - Being indexed says nothing about quality, and nobody can pay for it. Ask for a review: https://www.anchorterminal.com/builders/#claiming