Best of · Models & inference
Best fine-tuning services for AI models
All 9 ranked fine-tuning services for AI models on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 9 ranked
- 1 agent-ready
- 6 hosted endpoints
- Updated 8 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Highest score overall
Amazon Bedrock model customisation BB
BB, 75.8/100 on the benchmark.
Also Axolotl, B, 64.8/100.
Maintenance & community
Axolotl B
88/100 on maintenance & community, against 45 for the overall leader.
Transparency & trust
88/100 on transparency & trust, against 83 for the overall leader.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Amazon Bedrock model customisation Amazon Web Services |
BB 75.8 | Teams already on AWS that want to tune Amazon Nova or Llama models, or run reinforcement fine-tuning with a Lambda reward function, and serve the result inside Bedrock. | Pay per use | hosted |
| 2 | Axolotl Axolotl AI |
B 64.8 | A team that wants a repeatable, config-driven fine-tune of an open model on its own or rented GPUs, including multi-GPU and multi-node runs. | Free · OSS | library |
| 3 | Vertex AI Gemini tuning Google Cloud |
B 64.2 | Teams already on Google Cloud who need to tune Gemini itself, especially with RL, and will serve it there. | Pay per use | hosted |
| 4 | Microsoft Foundry fine-tuning (Azure OpenAI) Microsoft Azure |
C 61.1 | Teams that must tune an OpenAI model, need Azure's compliance and regional controls, and will serve the result on Azure. | Pay per use | hosted |
| 5 | Fireworks AI Fine-tuning Fireworks AI |
C 59 | Teams that want managed SFT, DPO or RFT on large open models and may later write a custom RL loop on the same platform. | Pay per use | hosted |
| 6 | Together AI Fine-tuning Together AI |
C 54.7 | Teams that want to tune a large open model, possibly with full fine-tuning, and take the weights away. | Pay per use | hosted |
| 7 | Unsloth Unsloth |
D 51.5 | One person or a small team tuning an open model on their own GPU and keeping the weights. | Free · OSS | library |
| 8 | Tinker Thinking Machines Lab |
D 51 | Researchers and teams writing custom post-training loops, especially RL, who want per-token billing and the weights at the end. | Pay per use | local |
| 9 | Nebius Token Factory fine-tuning Nebius |
D 47.7 | Teams that want supervised LoRA or full fine-tuning of a wide list of open models, up to Qwen3 Coder 480B and DeepSeek, through OpenAI-style calls, with EU storage and the weights to take away. | Pay per use | hosted |
How to choose
- Weight export and licenceCheck whether the tuned weights can be downloaded and under what licence, because an agent that must run offline or on your own hardware needs them.
- Tunable base models and methodsCheck which base models can be tuned and whether your chosen method is supported for each model, since method support can differ between them.
- Training price by methodCheck the training price per token for supervised and reinforcement runs separately, because the reinforcement price can differ from the supervised price on the same data.
- Serving and idle costsCheck the serving price, idle deletion and checkpoint storage fees, because an agent that calls a tuned model only occasionally can end up paying for idle time.
How the benchmark tests this category. The same small dataset used to tune a comparable open model on each service, then served. We check the job flow, how long training takes, whether the weights can leave, and the training and serving cost.
Each one in detail
Amazon Bedrock model customisation
BB 75.8/100Managed supervised fine-tuning, reinforcement fine-tuning and distillation of Amazon Nova, Meta Llama and selected open-weight models on Amazon Bedrock, run as asynchronous jobs through the Bedrock control-plane API, the AWS SDKs and CLI, or OpenAI-compatible endpoints.
Verdict Job creation takes an idempotency token, job lists filter and paginate, and the Service Terms give the customer exclusive use of a tuned model. Jobs run in two US Regions only, weights can't be exported, and the newest customisation API change found dates from 28 May 2026.
Choose it for Teams already on AWS that want to tune Amazon Nova or Llama models, or run reinforcement fine-tuning with a Lambda reward function, and serve the result inside Bedrock.
Strengths
CreateModelCustomizationJobaccepts aclientRequestToken, so a repeated create after a timeout doesn't start a second jobListModelCustomizationJobsfilters by status, name and creation time, sorts, and pages withmaxResultsup to 1,000 andnextToken- Service Terms section 50.12.4 gives the customer exclusive use of a customised model and bars third-party model providers from accessing it
Weaknesses
- Fine-tuning runs in us-east-1 and us-west-2 only, and each base model in one of them (two for the Titan models)
- No weight export was found in the reviewed documentation, and Service Terms section 50.11 forbids extracting model weights
- Setup needs an IAM service role,
iam:PassRoleand S3 buckets for input and output before the first job
Price Pay per useAuth OAuth or keyx402 nohosted
Axolotl
B 64.8/100Open-source command-line tool and Python package for fine-tuning open language models from one YAML config, covering LoRA, QLoRA, full fine-tuning, preference tuning and GRPO on the owner's GPUs.
Verdict Axolotl runs a whole fine-tuning job from one YAML file and ships a JSON Schema of its config plus bundled agent docs. It is 0.x software with telemetry on by default, no terms or privacy policy, and the owner supplies the GPU.
Choose it for A team that wants a repeatable, config-driven fine-tune of an open model on its own or rented GPUs, including multi-GPU and multi-node runs.
Strengths
- Apache-2.0, free, and the weights stay on the owner's hardware
axolotl config-schemaprints the full config as JSON Schema, andaxolotl agent-docsprints bundled Markdown references by topic- SFT, LoRA, QLoRA, DPO, IPO, KTO, ORPO, GRPO and reward modelling from one config format
Weaknesses
- Telemetry to PostHog is on by default and delays training start by 10 seconds until the variable is set either way
- No terms of service, privacy policy, legal entity or security.txt found on axolotl.ai
- Version 0.20.0, with removals in minor releases (FSDP1 in 0.20.0,
relora_stepsrenamed in 0.17.0 with no shim)
Price Free · OSSAuth Nonex402 nolibrary
Full assessment · Against #1, Amazon Bedrock model customisation
Vertex AI Gemini tuning
B 64.2/100Supervised, preference and reinforcement tuning of Gemini, plus supervised tuning of Gemma, Llama and Qwen, on Google Cloud's Gemini Enterprise Agent Platform (the platform formerly called Vertex AI).
Verdict Supervised, preference and reinforcement tuning of Gemini, plus supervised tuning of Gemma, Llama and Qwen. No weight export. The tuned model exists only as a Google Cloud endpoint.
Choose it for Teams already on Google Cloud who need to tune Gemini itself, especially with RL, and will serve it there.
Strengths
- Supervised, preference and reinforcement tuning of Gemini, plus supervised tuning of Gemma, Llama and Qwen
- No Vertex or Gemini incidents on the Google Cloud status dashboard from July to September 2026
- Public proto for GenAiTuningService with filter and pagination on job lists, and docs pages served as Markdown at
.md.txt
Weaknesses
- No weight export. The tuned model exists only as a Google Cloud endpoint
- Tuned Gemini 3 inference costs 1.5x the base model for as long as you serve it
- Setup needs a project, billing, IAM and a Cloud Storage bucket before the first job
Price Pay per useAuth OAuthx402 nohosted
Full assessment · Against #1, Amazon Bedrock model customisation
Microsoft Foundry fine-tuning (Azure OpenAI)
C 61.1/100Azure's managed service for supervised, preference and reinforcement fine-tuning of supported OpenAI and open-weight models.
Verdict SFT, DPO and RFT on GPT-4.1 and o4-mini through the OpenAI-shaped /openai/v1 API. No weight export; checkpoints copy only between Azure resources.
Choose it for Teams that must tune an OpenAI model, need Azure's compliance and regional controls, and will serve the result on Azure.
Strengths
- SFT, DPO and RFT on GPT-4.1 and o4-mini through the OpenAI-shaped /openai/v1 API
- Retirement policy with 60 days' notice and published training and deployment retirement dates per tunable model
- Entra ID with RBAC, Azure Monitor logs and an activity log for every customer
Weaknesses
- No weight export; checkpoints copy only between Azure resources
- $1.70 an hour hosting on Standard deployments, and deletion after 15 idle days
- GPT-4.1 training at $25 per 1M tokens globally, and no free tier without a card
Price Pay per useAuth OAuth or keyx402 nohosted
Full assessment · Against #1, Amazon Bedrock model customisation
Fireworks AI Fine-tuning
C 59/100Managed supervised, preference and reinforcement fine-tuning for open models, with a training API for custom workflows.
Verdict SFT, DPO, ORPO and RFT as managed jobs, plus a serverless Training API that is generally available. Tuned LoRAs only deploy to on-demand GPUs at $8 an hour and up, never to serverless.
Choose it for Teams that want managed SFT, DPO or RFT on large open models and may later write a custom RL loop on the same platform.
Strengths
- SFT, DPO, ORPO and RFT as managed jobs, plus a serverless Training API that is generally available
- LoRA SFT from $0.50 per 1M training tokens up to 16B parameters, with serving at base-model prices
- List endpoints take readMask, pageSize up to 200, AIP-160 filters and orderBy
Weaknesses
- Tuned LoRAs only deploy to on-demand GPUs at $8 an hour and up, never to serverless
- No training without a payment method; the $1 sign-up credit buys inference only
- The status page has no training component; a 46-hour cloud provider incident in August hit dedicated deployments
Price Pay per useAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Bedrock model customisation
Together AI Fine-tuning
C 54.7/100Managed LoRA and full fine-tuning, supervised or DPO, on about 30 open models from Qwen3.5 0.8B to Kimi K2.7, billed per training token with a $4 minimum.
Verdict 31 tunable base models, 11 or 12 of them with full fine-tuning as well as LoRA. Fine-tuned models don't run serverless; dedicated endpoints start at $5.49 an hour.
Choose it for Teams that want to tune a large open model, possibly with full fine-tuning, and take the weights away.
Strengths
- 31 tunable base models, 11 or 12 of them with full fine-tuning as well as LoRA
- GET /v1/finetune/download returns merged weights or the adapter, at any saved checkpoint
- POST /v1/fine-tunes/estimate-price quotes a job before it runs
Weaknesses
- Fine-tuned models don't run serverless; dedicated endpoints start at $5.49 an hour
- No free trial, a $5 prepaid purchase before the first call, and job minimums up to $60
- The status page covers serverless models only, and no fine-tuning rate limits are published
Price Pay per useAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Bedrock model customisation
Unsloth
D 51.5/100Open-source library, web UI (Studio) and desktop app for LoRA, QLoRA, full fine-tuning and RL (GRPO, DPO, ORPO) of open models on your own GPU, from 3 GB of VRAM.
Verdict The Apache-2.0 core runs on customer hardware and keeps model weights there. Users supply and pay for the GPU.
Choose it for One person or a small team tuning an open model on their own GPU and keeping the weights.
Strengths
- Free and open source, Apache-2.0 core, with the weights staying on your hardware
- LoRA, QLoRA, full fine-tuning, GRPO, DPO and ORPO from one package
- Exports adapters, merged 16-bit weights and GGUF for vLLM, Ollama or llama.cpp
Weaknesses
- Not a hosted service; you bring and pay for the GPU
- Studio is AGPL-3.0, and its server-side tools are on by default when exposed
- 792 open issues and 472 open pull requests
Price Free · OSSAuth Nonex402 nolibrary
Full assessment · Against #1, Amazon Bedrock model customisation
Tinker
D 51/100Thinking Machines Lab's API for model training.
Verdict Full control of the training loop with the GPUs abstracted away, plus recipes for SFT, DPO, RL and distillation. LoRA only; no full-parameter training.
Choose it for Researchers and teams writing custom post-training loops, especially RL, who want per-token billing and the weights at the end.
Strengths
- Full control of the training loop with the GPUs abstracted away, plus recipes for SFT, DPO, RL and distillation
- Checkpoints download and merge into Hugging Face safetensors, so the weights can leave
- Per-token billing with machine-readable prices in models.json
Weaknesses
- LoRA only; no full-parameter training
- Python SDK only, with no REST reference or OpenAPI
- No terms of service, status page or SLA found
Price Pay per useAuth API keyx402 nolocal
Full assessment · Against #1, Amazon Bedrock model customisation
Nebius Token Factory fine-tuning
D 47.7/100Nebius Token Factory runs supervised fine-tuning jobs on open models such as Llama, Qwen, gpt-oss, Gemma and DeepSeek through an OpenAI-compatible REST API, with LoRA or full weights and downloadable checkpoints.
Verdict Supervised fine-tuning on 49 open base models through OpenAI-style /v1/fine_tuning/jobs calls, with LoRA or full weights and every checkpoint file downloadable. No fine-tuning price was found outside the script-drawn console, and the docs say tuned models deploy only to dedicated endpoints, with custom weights in beta on request.
Choose it for Teams that want supervised LoRA or full fine-tuning of a wide list of open models, up to Qwen3 Coder 480B and DeepSeek, through OpenAI-style calls, with EU storage and the weights to take away.
Strengths
- 49 base models listed, 42 with LoRA and full fine-tuning and 7 with full fine-tuning only, at context lengths from 8,192 to 131,072 tokens
- Checkpoint files download through
GET /v1/files/{file_id}/content, and anhfintegration pushes the result to a Hugging Face repository - A public OpenAPI 3.1 file covers the fine-tuning, files, datasets and operations paths, with ranges on every hyperparameter
Weaknesses
- No fine-tuning price found in the docs or the public catalogue JSON. The price page is a script-drawn console page that robots.txt disallows
- The models page says deployment is by dedicated endpoints only, and custom model weights are in beta and available on request
- A bank card is mandatory at onboarding, so the $1 trial credit (30 days) is not a card-free trial
Price Pay per useAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Bedrock model customisation
Head to head
- Amazon Bedrock model customisation vs Axolotl BB 75.8 vs B 64.8
- Amazon Bedrock model customisation vs Vertex AI Gemini tuning BB 75.8 vs B 64.2
- Amazon Bedrock model customisation vs Microsoft Foundry fine-tuning (Azure OpenAI) BB 75.8 vs C 61.1
- Amazon Bedrock model customisation vs Fireworks AI Fine-tuning BB 75.8 vs C 59
- Axolotl vs Vertex AI Gemini tuning B 64.8 vs B 64.2
- Axolotl vs Microsoft Foundry fine-tuning (Azure OpenAI) B 64.8 vs C 61.1
- Axolotl vs Fireworks AI Fine-tuning B 64.8 vs C 59
- Microsoft Foundry fine-tuning (Azure OpenAI) vs Vertex AI Gemini tuning C 61.1 vs B 64.2
- Fireworks AI Fine-tuning vs Vertex AI Gemini tuning C 59 vs B 64.2
- Microsoft Foundry fine-tuning (Azure OpenAI) vs Fireworks AI Fine-tuning C 61.1 vs C 59
Questions
What are the highest-rated fine-tuning services for AI models for AI agents?
Amazon Bedrock model customisation has the highest benchmark score of the 9 ranked fine-tuning services for AI models, 75.8 (BB). Axolotl is second with 64.8 (B).
How many fine-tuning services for AI models are agent-ready?
1 of the 9 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which fine-tuning services for AI models accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 36 head-to-head comparisons · Best tools in every category