Category · Models & inference
Fine-tuning services for AI models
Services that train a model on your examples, by supervised, preference or reinforcement fine-tuning, and serve the result. Compared on which base models you can tune, the methods, price per training token, whether you get the weights and what serving the result costs.
Capability keys finetune.sft · finetune.preference · finetune.rl · finetune.lora · finetune.export · All tools
letme.dev/finetune.sft picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 190 | Vertex AI Gemini tuningGoogle Cloud · HTTP API | Fine-tuning | B | 64.2 | 2.5 (2) | n/a | n/a | Pay per use | OAuth | Hosted | ||
|
Supervised, preference and reinforcement tuning of Gemini, plus supervised tuning of Gemma, Llama and Qwen, on Google Cloud's Gemini Enterprise Agent Platform (the platform formerly called Vertex AI). Top strength Supervised, preference and reinforcement tuning of Gemini, plus supervised tuning of Gemma, Llama and Qwen Top weakness No weight export. The tuned model exists only as a Google Cloud endpoint |
||||||||||||
| 228 | Microsoft Foundry fine-tuning (Azure OpenAI)Microsoft Azure · HTTP API | Fine-tuning | C | 61.4 | 3.5 (2) | n/a | n/a | Pay per use | OAuth or key | Hosted | ||
|
Azure's managed service for supervised, preference and reinforcement fine-tuning of supported OpenAI and open-weight models. Top strength SFT, DPO and RFT on GPT-4.1 and o4-mini through the OpenAI-shaped /openai/v1 API Top weakness No weight export; checkpoints copy only between Azure resources |
||||||||||||
| 269 | Fireworks AI Fine-tuningFireworks AI · HTTP API | Fine-tuning | C | 59.2 | 2.5 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Managed supervised, preference and reinforcement fine-tuning for open models, with a training API for custom workflows. Top strength SFT, DPO, ORPO and RFT as managed jobs, plus a serverless Training API that is generally available Top weakness Tuned LoRAs only deploy to on-demand GPUs at $8 an hour and up, never to serverless |
||||||||||||
| 319 | Together AI Fine-tuningTogether AI · HTTP API | Fine-tuning | C | 54.9 | 3.0 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Managed LoRA and full fine-tuning, supervised or DPO, on about 30 open models from Qwen3.5 0.8B to Kimi K2.7, billed per training token with a $4 minimum. Top strength 31 tunable base models, 11 or 12 of them with full fine-tuning as well as LoRA Top weakness Fine-tuned models don't run serverless; dedicated endpoints start at $5.49 an hour |
||||||||||||
| 347 | UnslothUnsloth · Agent framework | Fine-tuning | D | 51.7 | 3.5 (2) | n/a | n/a | Free · OSS | None | Library | ||
|
Open-source library, web UI (Studio) and desktop app for LoRA, QLoRA, full fine-tuning and RL (GRPO, DPO, ORPO) of open models on your own GPU, from 3 GB of VRAM. Top strength Free and open source, Apache-2.0 core, with the weights staying on your hardware Top weakness Not a hosted service; you bring and pay for the GPU |
||||||||||||
| 354 | TinkerThinking Machines Lab · SDK + MCP | Fine-tuning | D | 51.2 | 3.5 (2) | n/a | n/a | Pay per use | API key | Local | ||
|
Thinking Machines Lab's API for model training. Top strength Full control of the training loop with the GPUs abstracted away, plus recipes for SFT, DPO, RL and distillation Top weakness LoRA only; no full-parameter training |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
How we test this category
The same small dataset used to tune a comparable open model on each service, then served. We check the job flow, how long training takes, whether the weights can leave, and the training and serving cost. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.




