# Novita AI (slim) > Novita AI is a hosted inference API for open-weight language models, with embeddings, reranking, image, video and speech models on the same key. It answers OpenAI-style calls at api.novita.ai with a Bearer key. - Full: https://www.anchorterminal.com/tools/novita-ai.md (~9,300 tokens) · this version ~1,980 tokens · JSON https://www.anchorterminal.com/tools/novita-ai.json · canonical https://www.anchorterminal.com/tools/novita-ai - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **D · 53.5/100 · rank #706 of 950 · #15 in Model APIs & inference · not agent-ready · confidence medium** Assessment: Per-token prices for about 140 language models are public, and each key can carry an expiry, a model allowlist and a source IP allowlist. No status page, SLA or security.txt was found, the per-model rate limit figures are drawn by script, and both vendor SDK repositories are archived while the docs still point to them. ## Facts - Kind: Model API · vendor: Novita AI · category: Model APIs & inference · legal entity: Novita AI · provenance 69/100 - Endpoint: `https://api.novita.ai/openai` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary service under the Novita AI Terms of Service. The agent skill and the MCP server are MIT - Probe metrics: not measured yet (probes haven't run) - Endpoints: OpenAI-style chat, completions, embeddings, models, files and batches under https://api.novita.ai/openai/v1, a rerank route, and separate task-based routes for image, video and audio models under https://api.novita.ai - Models on 9 October 2026: The pricing page lists 143 priced model rows across language, embedding, image, video and audio, and the home page says 200+ models. The live model list was not called - Rate limits: RPM and TPM per model, by five account tiers set by the highest monthly top-up in the last three calendar months (under $50, $50, $500, $3,000 and $10,000). The figures are drawn by script and were not read - Errors: 429 `RATE_LIMIT_EXCEEDED` or `TOKEN_LIMIT_EXCEEDED`, 403 `NOT_ENOUGH_BALANCE` or `model_access_denied`, 404 `MODEL_NOT_FOUND`, 503 `SERVICE_NOT_AVAILABLE`. The docs advise exponential backoff and client timeouts over 60 seconds - Billing by status: 200 and 499 are charged. 400, 401, 403, 429, 500, 503 and 504 are not - Batch: OpenAI-style files and batches for `/v1/chat/completions` and `/v1/completions`, one model per file, a fixed 24-hour window, input files kept 15 days, and an introductory 50 per cent discount - Prompt caching: Automatic on supported models with no request change. Cache-read prices are on the pricing page and `cached_tokens` is returned in usage. The FAQ puts cache reads at about a tenth of the input price - Structured output: The chat reference documents `json_object` and `json_schema` with `strict`. The LLM FAQ says only `json_object` is supported - Tool calling: `tools` with function definitions and a `strict` flag. The list of supporting models is drawn by script - Credentials: Bearer keys with the `sk_` prefix, 10 an account, expiry of 24 hours, 30 days, 90 days or permanent fixed at creation, deleted in the console. Per-key model access and IP access policies by API. OAuth with PKCE and scopes `api` and `balance:read` - Data use in the terms: Zero data retention and no use of content to train Novita's models or improve the services by default, except as law requires or as needed to run or support the service. Automated safety screening is allowed - Retirement notice: Notices appear in the changelog with a UTC time and a replacement model. The one read gave 11 days. No written notice period was found - SDKs: The docs use the OpenAI client with a changed base URL. npm `novita-sdk` 3.1.3 and Python `novita-client` come from repositories that are now archived - Agent files: `llms.txt`, `llms-full.txt`, `auth.md`, an API catalogue, an MCP server card, an agent card and a skill repository at version 1.4.0 - Other products on the same key: Agent Sandbox, GPU instances, serverless GPUs, dedicated endpoints, and resold Exa and Tavily search calls. Not graded in this listing - Models ($/1M in/out): `deepseek/deepseek-v4.1-flash` $0.30/$1.20; `qwen/qwen3.8-flash` $0.15/$0.47; `openai/gpt-oss-120b` $0.05/$0.25 - 2026-10-09 Shutdown: `deepseek/deepseek-v4-flash` retired at 16:00 UTC, replaced by `deepseek/deepseek-v4.1-flash` - Scores: Reliability 28, Performance pending, Schema & documentation 62, Agent ergonomics 72, Security & auth 65, Payments & pricing 35, Task success pending, Maintenance & community 62, Transparency & trust 56 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, No machine payment protocol was found in the docs index, the site map, the pricing page or the skill repository (0). · Maintenance & community, Model reading. · Transparency & trust, Closed service. - Sources: 37, open questions: 14, both in the full twin - Capabilities: inference.llm, inference.open-weights, embed.text, rerank, image.generate, video.generate, speech.tts, compute.batch - JSON: https://www.anchorterminal.com/api/v1/tools/novita-ai.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/novita-ai.svg` or a link to https://www.anchorterminal.com/tools/novita-ai from a page on novita.ai or one of its subdomains, or the README of github.com/novitalabs/novita-skills, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Set the OpenAI client base URL to `https://api.novita.ai/openai` and send the key as `Authorization: Bearer`. Keys start with `sk_` 2. Do not close a connection to stop generation. A request that reached the model is billed in full under status 499, so cap output with `max_tokens` 3. On 429, check whether the code is `RATE_LIMIT_EXCEEDED` or `TOKEN_LIMIT_EXCEEDED` and back off exponentially. No `Retry-After` header is documented 4. Read the changelog for retirement notices before pinning a model id. One retirement in October 2026 came with 11 days' notice 5. Ask the account owner for a key with an expiry, a model access policy and an IP allowlist. Keys are created and deleted only in the console ## Connect ```bash pip install openai # or: npm install openai ``` ```bash curl "https://api.novita.ai/openai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $NOVITA_API_KEY" \ -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"Hello"}]}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Cloudflare Workers AI | B | 68.3 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.tts, compute.batch | https://www.anchorterminal.com/tools/cloudflare-workers-ai.min.md | | DeepInfra | B | 63 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.tts, compute.batch | https://www.anchorterminal.com/tools/deepinfra.min.md | | SiliconFlow | D | 46.7 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, video.generate, speech.tts | https://www.anchorterminal.com/tools/siliconflow.min.md | | LocalAI | B | 68 | inference.open-weights, embed.text, rerank, speech.tts, image.generate, video.generate | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.open-weights, embed.text, rerank, speech.tts, image.generate | https://www.anchorterminal.com/tools/lemonade.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)