Category · Models & inference

Fine-tuning services for AI models

Services that train a model on your examples, by supervised, preference or reinforcement fine-tuning, and serve the result. Compared on which base models you can tune, the methods, price per training token, whether you get the weights and what serving the result costs.

Capability keys finetune.sft · finetune.preference · finetune.rl · finetune.lora · finetune.export · All tools

letme.dev/finetune.sft picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

6listings graded
0agent-ready (BB+)
12desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
6 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
190 Vertex AI Gemini tuningGoogle Cloud · HTTP API Fine-tuning B 64.2 2.5 (2) Pay per use
228 Microsoft Foundry fine-tuning (Azure OpenAI)Microsoft Azure · HTTP API Fine-tuning C 61.4 3.5 (2) Pay per use
269 Fireworks AI Fine-tuningFireworks AI · HTTP API Fine-tuning C 59.2 2.5 (2) Pay per use
319 Together AI Fine-tuningTogether AI · HTTP API Fine-tuning C 54.9 3.0 (2) Pay per use
347 UnslothUnsloth · Agent framework Fine-tuning D 51.7 3.5 (2) Free · OSS
354 TinkerThinking Machines Lab · SDK + MCP Fine-tuning D 51.2 3.5 (2) Pay per use

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

How we test this category

The same small dataset used to tune a comparable open model on each service, then served. We check the job flow, how long training takes, whether the weights can leave, and the training and serving cost. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.