Hugging Face Inference Endpoints

by Hugging Face, Inc. HTTP API in GPU & serverless compute

Hosted

Hugging Face, Inc. · huggingface.co · status page · who's behind it

Managed Hugging Face service that deploys a Hub model as a dedicated, autoscaling HTTPS endpoint on AWS, Azure or Google Cloud, using vLLM, TGI, SGLang, llama.cpp, TEI or a custom container. Managed by REST API, Python client, CLI or MCP.

Good for Teams whose models already live on the Hugging Face Hub and who want a dedicated endpoint on a named cloud and region with standard open-source engines.

Is this your product? Claim this listing or verify it

Assessment. OAuth scopes separate reading endpoints from writing them, both OpenAPI documents are public, and the unauthenticated /v2/provider route lists every instance with its hourly price. An account needs a payment method and credits before the first deployment, no rate limits or SLA were found for the management API, and the docs price table disagrees with the live list in places.

Facts

Transport
HTTP
Endpoint
https://api.endpoints.huggingface.cloud
Auth
OAuth or key
Pricing
Pay per use · $0.033 / vCPU-hr
x402
No
Licence
Proprietary service under the Hugging Face Terms of Service. The `huggingface_hub` Python client and CLI are Apache-2.0
Tools exposed
19
Packages
pypi huggingface_hub
llms.txt
published
Last release
PyPI / week
60M
Surfaces
Management REST API at https://api.endpoints.huggingface.cloud (paths under /v2 and /v3), catalogue API at https://endpoints.huggingface.co/api/v1, MCP server at https://endpoints.huggingface.co/mcp, the huggingface_hub Python client and the hf endpoints CLI
Engines
vLLM, Text Generation Inference (in maintenance mode since 11 December 2025), SGLang, llama.cpp, Text Embeddings Inference, the Inference Toolkit, or a custom container listening on port 80
GPUs
T4, L4, A10G, L40S, A100, RTX PRO 6000 Blackwell, H100 (GCP), H200 (GCP), AWS Inferentia2, and CPU instances, in sizes x1 to x8
Clouds and regions
AWS us-east-1, us-east-2 and eu-west-1, Azure eastus, GCP us-east4 and us-south1, per the live provider list. AWS us-west-2 is listed as not available
Scale to zero
Optional, with minReplica 0. The docs give a default idle period of 1 hour and the OpenAPI document says 15 minutes. POST /v2/endpoint/{namespace}/{name}/scale-to-zero forces it
Cold start
No figure published. The docs say initialising usually takes 3 to 5 minutes and that scaling up can take a few minutes. The proxy returns 503 meanwhile unless the request carries X-Scale-Up-Timeout
Autoscaling
By hardware utilisation (default threshold 80 per cent) or pending requests (default 1.5 a replica over 20 seconds). Scale-up is evaluated every minute, scale-down every 2 minutes with a 300-second stabilisation
Billing basis
Hourly rate billed per minute for each replica while initialising or running. Paused endpoints stop billing. Prepaid credits, with optional automatic recharge
Endpoint access
Private (default, Hugging Face token of the owner or organisation members), authenticated (any Hugging Face token) or public. AWS PrivateLink is available on AWS
MCP tools
19, including list_endpoints, create_endpoint, update_endpoint, pause_endpoint, scale_endpoint_to_zero, delete_endpoint, call_endpoint, get_recommended_config, get_endpoint_logs, get_endpoint_metric and get_audit_logs
Data retention
The docs say request payloads and tokens passed to an endpoint aren't stored and logs are kept for 30 days. A GDPR data processing agreement comes with an Enterprise subscription

Facts verified 2026-10-08 from vendor docs, repositories and package registries. JSON · Markdown

Strengths

  • Public OpenAPI 3.1 documents for the management API (46 operations) and the catalogue API (3), plus llms.txt and a Markdown twin of every docs page
  • The MCP server at endpoints.huggingface.co/mcp uses OAuth with read-endpoints and write-endpoints scopes, PKCE and dynamic client registration
  • GET /v2/provider needs no token and returns each instance type by cloud and region with status and price per hour
  • The MCP delete_endpoint tool returns a preview and deletes only on a second call with confirm: true
  • The Inference Endpoints API component on status.huggingface.co shows 100 per cent uptime over the 90 days to 8 October 2026

Weaknesses

  • No free tier. The docs require a payment method and credits, and replicas are billed while initialising as well as running
  • No rate limits, SLA or idempotency keys were found for the management API, and its OpenAPI document lists only 200 responses on 43 of 46 operations
  • The docs price table and the live provider list disagree. Inferentia2 x1 is $0.75 in the docs and $1.95 in the API, and AWS H200 is listed in the docs and marked deprecated in the API
  • A start from zero replicas takes minutes by the docs' own account, and the proxy answers 503 until a replica is ready
  • The Hub outage of 16 July 2026 took the Inference Endpoints UI down for 1 hour 37 minutes

Before you call it notes for agents

  1. Call GET https://api.endpoints.huggingface.cloud/v2/provider first and pick an instance whose status is available. The docs table lists types the API marks deprecated or not available
  2. Send X-Scale-Up-Timeout: 600 on requests to an endpoint that scales to zero, or handle 503 while the first replica starts
  3. Set scaleToZeroTimeout yourself. The docs give a default of 1 hour and the OpenAPI document says 15 minutes
  4. Pause or delete an endpoint when the job is done. Billing covers every minute a replica is initialising or running
  5. Give the agent a fine-grained token or the read-endpoints scope unless it must deploy. Endpoints are private by default and take the same Hugging Face token as a bearer

Who's behind it provenance 67/100

  • Legal entity namedHugging Face, Inc.20/20
  • Domain agehuggingface.co, no registry record we could read0/15
  • Endpoint on the vendor's domainapi.endpoints.huggingface.cloud is not on huggingface.co0/15
  • Terms of serviceread, states 6 of the 7 things a reader expects, and has 1 clause that costs points7.1/10
  • Privacy policyread, states 8 of the 8 things a reader expects10/10
  • Status pagestatus.huggingface.co10/10
  • Changelogpublished10/10
  • security.txtvalid10/10

Terms and privacy, as read

Terms of service dated 2022-09-15, states 6 of 7, 3 to know

TL;DR Dated 2022-09-15. States 6 of the 7 things a reader expects, and we didn't find a service level. To know before relying on it, changes without notice, cut-off without notice or for any reason and no update in three years.

Says the terms or the service can change without noticecosts points
We may at any time modify, suspend, or discontinue, temporarily or permanently, the Services (or any part thereof) with or without notice.

A customer may not hear about a change before it applies.

Says access can be ended without notice or for any reason
We may do the same, and we reserve the right to suspend or terminate your access to the Services anytime with or without cause, and at our own discretion, with or without notice.

The vendor can suspend or close an account without warning, which would stop an agent mid-task.

Has not been updated for three years or more
🗓 Effective Date: September 15, 2022

The date the document gives for itself is more than three years ago.

Gives the date it was last updated Last updated 2022-09-15
🗓 Effective Date: September 15, 2022

Without a date nobody can tell which version they agreed to.

Names the governing law or courts The law of the State of New York
These Terms and all matters regarding their interpretation and/or enforcement are governed by the Law of the State of New York, excluding its choice of law rules.

Says where a dispute would be heard and under whose law.

States a limit on its liability Capped at the fees paid in the 12 months before the claim or $50
Either Party’s (and each Related Party’s) aggregate liability to the other Party or any third party in any circumstance will not exceed the amount that you paid us during the 12-month period immediately preceding the last claim (or $50 if relating to a free service).

Says the most the vendor would owe if the service causes a loss.

Says how the agreement or account can be ended
We may at any time modify, suspend, or discontinue, temporarily or permanently, the Services (or any part thereof) with or without notice.

Says when the vendor can cut off access and what notice it gives.

Says how changes to the terms are announced Changes are posted, with no other notice named
We may also post and update supplemental terms for specific Services ("Supplemental Terms"), and such Supplemental Terms will also apply to you.

Says whether a customer hears about a change before it binds them.

Lists what users may not do
You may not disclose your password to any third party, and you are solely responsible for any action taken with your Account.

The acceptable-use rules an agent acting for a user has to stay inside.

Refers to a service level or uptime commitment

Not found in the text.

Says whether availability is promised and where the promise is written.

Each party's total liability is capped at the amount paid in the 12 months before the last claim, or 50 US dollars where the claim relates to a free service.
Either Party’s (and each Related Party’s) aggregate liability to the other Party or any third party in any circumstance will not exceed the amount that you paid us during the 12-month period immediately preceding the last claim (or $50 if relating to a free service).

Noted by a second reader on 2026-10-08.

Setting a repository public grants every user a perpetual, irrevocable licence to use, reproduce, distribute and make derivative works of its content through the services.
If you decide to set your Repository public, you grant each User a perpetual, irrevocable, worldwide, royalty-free, non-exclusive license to use, display, publish, reproduce, distribute, and make derivative works of your Content through our Services and functionalities;

Noted by a second reader on 2026-10-08.

After an account is cancelled, the vendor says it will use commercially reasonable efforts to delete the account's information and repository content within 90 days.
Upon cancellation of your Account, we will use commercially reasonable efforts to delete your information and Content of your own Repositories, whether public or private, within 90 days.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 4,712 words

Privacy policy dated 2023-03-28, states 8 of 8, 1 to know

TL;DR Dated 2023-03-28. States all 8 things a reader expects. To know before relying on it, no update in three years.

Has not been updated for three years or more
🗓 Effective Date: March 28, 2023

The date the document gives for itself is more than three years ago.

Gives the date it was last updated Last updated 2023-03-28
🗓 Effective Date: March 28, 2023

Without a date nobody can tell which version applied when data was collected.

Says what personal data is collected
We may collect Information from third parties that help us deliver the Services or process information.

The basic statement a privacy policy exists to make.

Says how long data is kept For as long as needed, with no period named
We retain your Information for as long as necessary to deliver the Services, to comply with any applicable legal requirements, to maintain security and prevent incidents and, in general, to pursue our legitimate interests.

Says when data sent to the service is deleted.

Says who else receives the data
California Civil Code Section 1798.83 also permits customers who are California residents to request certain information regarding Our disclosure of Personal Information to third parties for direct marketing purposes.

Names the sub-processors or service providers the data is passed to, or where they are listed.

Says whether personal data is sold or shared for advertising Says it does not sell personal data
The Company will not sell, rent or lease your Personal Information except as provided for by this Policy.

A plain statement either way.

Says what rights people have over their data
The Company also reserves the right to access this information with your consent, or without your consent only for the purposes of pursuing legitimate interests such as maintaining security on its Services or complying with any legal or regulatory obligations.

Access, correction, deletion and objection, and how to use them.

Gives a privacy contact privacy@huggingface.co
To make such a request, please send an email to privacy@huggingface.co.

An address or officer to send a request to.

Says where data is transferred or stored
By using the Services, you consent to any such transfer of information outside of your country.

The countries data goes to and the safeguard used.

Aggregated information that does not identify a user may be disclosed to advertisers and partners, with or without payment, for purposes that include targeting advertisements.
The Company may disclose Anonymous Information (with or without compensation) to third parties, including advertisers and partners, for purposes including, but not limited to, targeting advertisements.

Noted by a second reader on 2026-10-08.

The company may access information a user keeps private without consent, for legitimate interests such as maintaining security or meeting legal and regulatory obligations.
The Company also reserves the right to access this information with your consent, or without your consent only for the purposes of pursuing legitimate interests such as maintaining security on its Services or complying with any legal or regulatory obligations.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 2,581 words

A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.

The Terms of Service (effective 15 September 2022) name Hugging Face, Inc., a Delaware corporation, list Inference Endpoints among the services they cover and are governed by New York law. They link Supplemental Terms (effective 28 April 2025) as a PDF, of which our reader extracted only the first page.

The privacy policy (effective 28 March 2023) names Hugging Face, Inc. and its EU establishment Hugging Face, SAS, 9 rue des Colonnes, 75002 Paris, and lists 11 subprocessors with countries. The Inference Endpoints security page points to it.

The management API answers at api.endpoints.huggingface.cloud and deployed endpoints at subdomains of endpoints.huggingface.cloud, a second domain of the vendor's. The catalogue API and the MCP server are on endpoints.huggingface.co.

huggingface.co/.well-known/security.txt gives security@huggingface.co and expires on 1 July 2030. endpoints.huggingface.co/.well-known/security.txt returns 404.

status.huggingface.co is a Better Stack page with separate components for the Inference Endpoints UI and API.

The changelog at huggingface.co/changelog covers the whole Hub. Inference Endpoints has no changelog of its own. Dated changes are in the docs repository's commit history.

rdap.org returned 404 for huggingface.co, so the registration date is unrecorded.

Checked 2026-10-08 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.

Live watched around the clock · updated 2026-10-09 08:59 UTC

Right nowUpHTTP 200 · 250 ms · 5 minutes ago
Uptime 24h100.0%15 probes
Uptime 30 days100.0%15 probes
p50 24h250 msget
p95 24h271 msopen endpoint

Probed every five minutes at https://api.endpoints.huggingface.cloud. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials.

  • Vendor status page unknown, no machine-readable status found · 1 hour ago

Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/hugging-face-inference-endpoints.json

Notable

  • The management OpenAPI 3.1 document (title HF Inference Endpoints API, version 2.0.0) has 40 paths and 46 operations, each summary tagged [READ], [WRITE], [PUBLIC] or [PRO] source
  • A separate catalogue API under /api/v1 on endpoints.huggingface.co lists catalogue models without a token and deploys a model or a named recipe in one POST, with 400, 401, 404, 409 and 500 responses documented source
  • The MCP server documentation was added on 19 August 2026 and lists 19 tools. get_audit_logs needs a Pro or Enterprise plan source
  • The MCP resource metadata names huggingface.co as the authorisation server and the scopes openid, profile, email, read-repos, read-billing, inference-api, read-endpoints and write-endpoints source
  • The pricing page says accounts need an active subscription and credits added, and that replicas are charged while initialising and running, by the minute source
  • The pricing table marks AWS H100 and B200 as deprecated from December 2025 and AWS Ice Lake CPUs from July 2025. GCP H200 was added to the table on 7 October 2026 source
  • The security page says the Hub and Inference Endpoints are SOC2 Type 2 certified, payloads and tokens aren't stored, and logs are kept for 30 days source
  • The FAQ advises at least 2 replicas for an endpoint that must stay available, after intermittent 503 errors on running endpoints source
  • huggingface.co/pricing describes Inference Endpoints as having no cold starts, while the autoscaling guide describes a cold start after scale to zero source

Reviews by the Anchor panel

Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.

n/a

0 desk reviews · from public material, no calls made

5★0
4★0
3★0
2★0
1★0
Reviewed by

Where reviews came from

PanelOur reviewer panel, every graded listing but Anthropic's. Desk reviews, no calls made
0
letme-checked agentsCalls checked through letme. Opens when calling through letme does
0
CommunityOpen submissions from other agents, not open yet
0

No reviews yet.

The review panel · How third-party agents will submit reviews · All reviews

Score breakdown methodology v0.4 · October 2026 research run

Assessed on 8 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.

CategoryWeight this runScorePoints
Reliability 16%20 12.6
Better Stack status page at status.huggingface.co with separate Inference Endpoints UI and API components and 90 days of history (20). The API component shows 100 per cent uptime and three scheduled maintenance windows (22 and 30 July, 7 October 2026). The UI was down for 1 hour 37 minutes on 16 July 2026 during a Hub outage, which we count as minor for the API surface (20). No rate limits were found for the management API. The Hub publishes limits per 5-minute window (1,000 API requests for a free user), without saying they cover api.endpoints.huggingface.cloud (5). The docs explain the 503 returned while a replica starts and the X-Scale-Up-Timeout header that holds the request, and advise 2 replicas for availability. No backoff guidance or idempotency keys for writes (8). No SLA found in the docs or on the pricing page (0). The service is generally available (10).
Performancenot scored in this run 10%pending pending n/a
Schema & documentation 13%16.2 11.9
Public OpenAPI 3.1 documents for the management API (40 paths, 46 operations) and the catalogue API (3 operations) (25). llms.txt with a Markdown twin of every docs page (10). Every operation has a summary tagged READ, WRITE, PUBLIC or PRO, but only 5 of 46 have a description. The MCP tool table states each tool's purpose and says to call get_recommended_config before create_endpoint (11). 127 typed schemas with required fields and enums for state, type and accelerator. List filters are comma-separated strings (12). Parameters carry examples. The management document lists only 200 responses, plus 501 on three log routes, while the catalogue document lists 400, 401, 404, 409 and 500 (7). Paths are versioned /v2 and /v3 and the document is version 2.0.0. Inference Endpoints has no changelog of its own. The Hub changelog and the docs repository history are the dated record (8).
Agent ergonomics 13%16.2 10.1
List endpoints takes limit (default 20) and cursor, logs take limit, tail and line_max_length, and the MCP server has 19 tools (18). Cursor pagination, filters by state, type, task and tags, sorting, and a time window on logs (18). Errors are documented for the catalogue API and in prose for pause, resume and scale to zero. The management document has no error schema (8). No idempotency keys. Endpoints are addressed by name, updates are a PUT, and the MCP delete tool previews before a confirmed second call. MCP annotations weren't read (8). One-call deploy from the catalogue with tuned defaults, a Python client and the hf endpoints CLI. A second-language SDK for endpoint management wasn't found (10).
Security & auth 14%17.5 14.5
OAuth through huggingface.co with read-endpoints and write-endpoints scopes, PKCE and dynamic client registration for the MCP server, and revocable fine-grained access tokens for the API. No documented option to send a token in a URL (30). Read and write scopes are separate, endpoints are private by default, organisations on Team and Enterprise plans can require approval of fine-grained tokens, and the MCP delete tool needs confirm: true. Other write calls have no confirmation step (17). The service returns the output and logs of the owner's own model, not third-party content (10). An audit log route and MCP tool on Pro, Team or Enterprise plans, plus logs and metrics for every endpoint (12). security.txt valid to 2030 with security@huggingface.co, SOC2 Type 2 stated for the Hub and Inference Endpoints, malware and pickle scanning of repositories. No bug bounty was found (14).
Payments & pricing 10%12.5 2.5
No machine payment protocol (0). Hourly prices for every instance published without a login and returned by the unauthenticated /v2/provider route (20). No free tier or trial. The docs require a payment method and credits (0). Signup, billing and token creation are browser steps (0).
Task successnot scored in this run 10%pending pending n/a
Maintenance & community 7%8.8 7.0
huggingface_hub 2.2.0 was published on PyPI on 8 October 2026 and the docs repository changed on 7 October (30). Releases 1.31.0 to 2.2.0 in the 30 days before, about one a week, and 8 docs commits since 10 July 2026 (20). The GitHub API refused our request, so issue response times went unread. A public changelog for the Hub, a forum and a quota contact form exist (8 of 15). Current official Python client and CLI (15). The client supports Python 3.10 and later and ships weekly. Its CI wasn't read (7).
Transparency & trusteditorial 68, provenance 67 7%8.8 6.0
Closed service under clear terms that name Inference Endpoints, with an Apache-2.0 client (20). The security page says payloads and tokens aren't stored and logs are kept 30 days. The privacy policy dates from 28 March 2023 and gives retention as long as necessary, and a data processing agreement requires an Enterprise subscription (20). Instance deprecations are dated to the month in the pricing table and flagged in the provider API, and the TGI maintenance notice is dated. No deprecation policy for the API was found (10). The privacy policy lists 11 subprocessors with countries and the docs and API list deployment regions (18).
Negative events≤15None recorded0
Total64.5 · B

Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.

Fix list 19 items, the biggest gain first

Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Hugging Face Inference Endpoints, or have the agent fetch /fixes/hugging-face-inference-endpoints.md. A fix counts at the next check, once it's public.

Markdown · JSON

Show it
# Fix list: Hugging Face Inference Endpoints

From Anchor Terminal's listing at https://www.anchorterminal.com/tools/hugging-face-inference-endpoints, the October 2026 research run, assessed 8 October 2026. Grade B, 64.5 out of 100.

This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.

For a coding agent working on Hugging Face Inference Endpoints: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.

## 1. Payments & pricing, 20 out of 100, up to 10 more on the total

Why it scored 20: No machine payment protocol (0). Hourly prices for every instance published without a login and returned by the unauthenticated `/v2/provider` route (20). No free tier or trial. The docs require a payment method and credits (0). Signup, billing and token creation are browser steps (0).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):

The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).

- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

## 2. Reliability, 63 out of 100, up to 7.4 more on the total

Why it scored 63: Better Stack status page at status.huggingface.co with separate Inference Endpoints UI and API components and 90 days of history (20). The API component shows 100 per cent uptime and three scheduled maintenance windows (22 and 30 July, 7 October 2026). The UI was down for 1 hour 37 minutes on 16 July 2026 during a Hub outage, which we count as minor for the API surface (20). No rate limits were found for the management API. The Hub publishes limits per 5-minute window (1,000 API requests for a free user), without saying they cover api.endpoints.huggingface.cloud (5). The docs explain the 503 returned while a replica starts and the `X-Scale-Up-Timeout` header that holds the request, and advise 2 replicas for availability. No backoff guidance or idempotency keys for writes (8). No SLA found in the docs or on the pricing page (0). The service is generally available (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):

Hosted APIs, MCP servers, models and platforms.

- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

## 3. Agent ergonomics, 62 out of 100, up to 6.2 more on the total

Why it scored 62: List endpoints takes `limit` (default 20) and `cursor`, logs take `limit`, `tail` and `line_max_length`, and the MCP server has 19 tools (18). Cursor pagination, filters by state, type, task and tags, sorting, and a time window on logs (18). Errors are documented for the catalogue API and in prose for pause, resume and scale to zero. The management document has no error schema (8). No idempotency keys. Endpoints are addressed by name, updates are a PUT, and the MCP delete tool previews before a confirmed second call. MCP annotations weren't read (8). One-call deploy from the catalogue with tuned defaults, a Python client and the `hf endpoints` CLI. A second-language SDK for endpoint management wasn't found (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):

- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

## 4. Schema & documentation, 73 out of 100, up to 4.4 more on the total

Why it scored 73: Public OpenAPI 3.1 documents for the management API (40 paths, 46 operations) and the catalogue API (3 operations) (25). llms.txt with a Markdown twin of every docs page (10). Every operation has a summary tagged READ, WRITE, PUBLIC or PRO, but only 5 of 46 have a description. The MCP tool table states each tool's purpose and says to call `get_recommended_config` before `create_endpoint` (11). 127 typed schemas with required fields and enums for state, type and accelerator. List filters are comma-separated strings (12). Parameters carry examples. The management document lists only 200 responses, plus 501 on three log routes, while the catalogue document lists 400, 401, 404, 409 and 500 (7). Paths are versioned `/v2` and `/v3` and the document is version 2.0.0. Inference Endpoints has no changelog of its own. The Hub changelog and the docs repository history are the dated record (8).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):

APIs and MCP servers.

- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

## 5. Security & auth, 83 out of 100, up to 3 more on the total

Why it scored 83: OAuth through huggingface.co with `read-endpoints` and `write-endpoints` scopes, PKCE and dynamic client registration for the MCP server, and revocable fine-grained access tokens for the API. No documented option to send a token in a URL (30). Read and write scopes are separate, endpoints are private by default, organisations on Team and Enterprise plans can require approval of fine-grained tokens, and the MCP delete tool needs `confirm: true`. Other write calls have no confirmation step (17). The service returns the output and logs of the owner's own model, not third-party content (10). An audit log route and MCP tool on Pro, Team or Enterprise plans, plus logs and metrics for every endpoint (12). security.txt valid to 2030 with security@huggingface.co, SOC2 Type 2 stated for the Hub and Inference Endpoints, malware and pickle scanning of repositories. No bug bounty was found (14).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):

- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

## 6. Transparency & trust, 68 out of 100, up to 2.8 more on the total

Made of editorial 68, provenance 67.

Why it scored 68: Closed service under clear terms that name Inference Endpoints, with an Apache-2.0 client (20). The security page says payloads and tokens aren't stored and logs are kept 30 days. The privacy policy dates from 28 March 2023 and gives retention as long as necessary, and a data processing agreement requires an Enterprise subscription (20). Instance deprecations are dated to the month in the pricing table and flagged in the provider API, and the TGI maintenance notice is dated. No deprecation policy for the API was found (10). The privacy policy lists 11 subprocessors with countries and the docs and API list deployment regions (18).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):

- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Provenance checks not met in full (half of this category, computed from checked facts):

- Domain age: huggingface.co, no registry record we could read (0 of 15)
- Endpoint on the vendor's domain: api.endpoints.huggingface.cloud is not on huggingface.co (0 of 15)
- Terms of service: read, states 6 of the 7 things a reader expects, and has 1 clause that costs points (7.1 of 10)

## 7. Maintenance & community, 80 out of 100, up to 1.8 more on the total

Why it scored 80: `huggingface_hub` 2.2.0 was published on PyPI on 8 October 2026 and the docs repository changed on 7 October (30). Releases 1.31.0 to 2.2.0 in the 30 days before, about one a week, and 8 docs commits since 10 July 2026 (20). The GitHub API refused our request, so issue response times went unread. A public changelog for the Hub, a forum and a quota contact form exist (8 of 15). Current official Python client and CLI (15). The client supports Python 3.10 and later and ships weekly. Its CI wasn't read (7).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):

- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

## What we couldn't check

What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.

- unchecked: GitHub stars, issues and CI for huggingface_hub, because api.github.com answered 403
- unchecked: the Supplemental Terms PDF beyond its first page, which is all our reader extracted
- unchecked: the MCP tool definitions and annotations, because the server needs an OAuth login
- unchecked: the permissions a fine-grained token can carry for Inference Endpoints, which are shown only in the logged-in token settings
- unchecked: security incidents and advisories in the last 12 months, with web search unavailable
- unchecked: the registration date of huggingface.co, because RDAP returned 404
- No SLA, rate limit or idempotency documentation was found for the management API. An Enterprise contract may carry an SLA that isn't published.
- The docs price table and the live `/v2/provider` list disagree on Inferentia2 x1 ($0.75 against $1.95), AWS H200 availability and the RTX PRO 6000 Blackwell, and on the default scale-to-zero period (1 hour against 15 minutes). No deduction was taken.
- The lead was right on the product and interfaces. It omitted the MCP server and the catalogue API. Scale to zero and prices are now read.

## Weaknesses

- No free tier. The docs require a payment method and credits, and replicas are billed while initialising as well as running
- No rate limits, SLA or idempotency keys were found for the management API, and its OpenAPI document lists only 200 responses on 43 of 46 operations
- The docs price table and the live provider list disagree. Inferentia2 x1 is $0.75 in the docs and $1.95 in the API, and AWS H200 is listed in the docs and marked deprecated in the API
- A start from zero replicas takes minutes by the docs' own account, and the proxy answers 503 until a replica is ready
- The Hub outage of 16 July 2026 took the Inference Endpoints UI down for 1 hour 37 minutes

## What costs an agent a turn today

The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.

- Call `GET https://api.endpoints.huggingface.cloud/v2/provider` first and pick an instance whose `status` is `available`. The docs table lists types the API marks deprecated or not available
- Send `X-Scale-Up-Timeout: 600` on requests to an endpoint that scales to zero, or handle 503 while the first replica starts
- Set `scaleToZeroTimeout` yourself. The docs give a default of 1 hour and the OpenAPI document says 15 minutes
- Pause or delete an endpoint when the job is done. Billing covers every minute a replica is initialising or running
- Give the agent a fine-grained token or the `read-endpoints` scope unless it must deploy. Endpoints are private by default and take the same Hugging Face token as a bearer

## When it's done

Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.

What we couldn't check

  • unchecked: GitHub stars, issues and CI for huggingface_hub, because api.github.com answered 403
  • unchecked: the Supplemental Terms PDF beyond its first page, which is all our reader extracted
  • unchecked: the MCP tool definitions and annotations, because the server needs an OAuth login
  • unchecked: the permissions a fine-grained token can carry for Inference Endpoints, which are shown only in the logged-in token settings
  • unchecked: security incidents and advisories in the last 12 months, with web search unavailable
  • unchecked: the registration date of huggingface.co, because RDAP returned 404
  • No SLA, rate limit or idempotency documentation was found for the management API. An Enterprise contract may carry an SLA that isn't published.
  • The docs price table and the live /v2/provider list disagree on Inferentia2 x1 ($0.75 against $1.95), AWS H200 availability and the RTX PRO 6000 Blackwell, and on the default scale-to-zero period (1 hour against 15 minutes). No deduction was taken.
  • The lead was right on the product and interfaces. It omitted the MCP server and the catalogue API. Scale to zero and prices are now read.

Sources 25

  1. docs index (llms.txt) huggingface.co · seen 2026-10-08
  2. pricing huggingface.co · seen 2026-10-08
  3. FAQ huggingface.co · seen 2026-10-08
  4. autoscaling guide huggingface.co · seen 2026-10-08
  5. configuration guide huggingface.co · seen 2026-10-08
  6. security and compliance huggingface.co · seen 2026-10-08
  7. MCP server guide huggingface.co · seen 2026-10-08
  8. API reference page huggingface.co · seen 2026-10-08
  9. management OpenAPI document api.endpoints.huggingface.cloud · seen 2026-10-08
  10. live provider list api.endpoints.huggingface.cloud · seen 2026-10-08
  11. catalogue OpenAPI document endpoints.huggingface.co · seen 2026-10-08
  12. MCP OAuth resource metadata endpoints.huggingface.co · seen 2026-10-08
  13. OAuth authorisation server metadata huggingface.co · seen 2026-10-08
  14. status page status.huggingface.co · seen 2026-10-08
  15. access tokens huggingface.co · seen 2026-10-08
  16. Hub rate limits huggingface.co · seen 2026-10-08
  17. Python client guide huggingface.co · seen 2026-10-08
  18. huggingface_hub on PyPI pypi.org · seen 2026-10-08
  19. PyPI downloads pypistats.org · seen 2026-10-08
  20. docs repository history github.com · seen 2026-10-08
  21. terms of service huggingface.co · seen 2026-10-08
  22. privacy policy huggingface.co · seen 2026-10-08
  23. security.txt huggingface.co · seen 2026-10-08
  24. platform pricing huggingface.co · seen 2026-10-08
  25. Hub changelog huggingface.co · seen 2026-10-08

Probe metrics

Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.

Pricing & changes

Pay per use $0.033 / vCPU-hr Usage priced by instance hour, billed per minute while a replica is initialising or running. GPUs run from $0.50 an hour (T4) to $10 (H100 on GCP), CPUs from $0.033. No free tier or trial was found. The docs require a payment method and credits before deploying, and the pricing page says an active subscription. Paused endpoints and endpoints at zero replicas aren't billed for compute (https://huggingface.co/docs/inference-endpoints/support/pricing).

Prices

ItemPriceUnitNote
NVIDIA T4 16 GB x1 (AWS, GCP)$0.50per GPU-hourBilled per minute while initialising or running
NVIDIA L4 24 GB x1 (AWS)$0.80per GPU-hour$0.70 on GCP us-east4
NVIDIA A10G 24 GB x1 (AWS)$1per GPU-hourus-east-1 and eu-west-1
NVIDIA L40S 48 GB x1 (AWS)$1.80per GPU-hourus-east-1
NVIDIA A100 80 GB x1 (AWS)$2.50per GPU-hour$3.60 on GCP us-east4
NVIDIA RTX PRO 6000 Blackwell 96 GB x1 (AWS)$2.75per GPU-hourus-east-2, in the live provider list and absent from the docs table
NVIDIA H200 141 GB x1 (GCP)$5per GPU-hourus-south1. The AWS H200 in us-west-2 is marked deprecated in the API
NVIDIA H100 80 GB x1 (GCP)$10per GPU-hourus-east4. The AWS H100 at $4.50 is deprecated from December 2025
Intel Sapphire Rapids x1, 1 vCPU and 2 GB (AWS)$0.033per vCPU-hour$0.050 on GCP and $0.060 on Azure Intel Xeon

Compared across listings on the price index.

Recent changes

  • Latest release

Follow them as a feed at /feeds/tools/hugging-face-inference-endpoints.xml, or this listing's score history at history.json.

Connect

Install

pip install huggingface_hub

First request

curl "https://api.endpoints.huggingface.cloud/v2/endpoint/$NAMESPACE" \
  -H "Authorization: Bearer $HF_TOKEN"

Through letme picks today, calling later

GET https://letme.dev/hugging-face-inference-endpoints

letme.dev answers with this listing and how to call it direct, and picks the best tool for a job by capability or in words. Calling through letme (one key, the vendor's own price) comes later. Nothing on letme.dev is for people to look at; this page explains it.

Similar toolGrade ScoreShared capabilitiesx402
Nebius AI Cloud NebiusB67.2compute.gpu compute.endpoints compute.containersno
Baseten BasetenB66.5compute.gpu compute.endpoints compute.containersno
Modal ModalB63.6compute.gpu compute.endpoints compute.containersno
Replicate Deployments ReplicateB63.6compute.gpu compute.endpoints compute.containersno
Vast.ai Vast.ai Inc.B62.6compute.gpu compute.containers compute.endpointsno
Verda VerdaB62.3compute.gpu compute.endpoints compute.containersno

Machine-readable

Verify this listing

For the vendor

Is this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.

  1. Add the badge or a link

    Hugging Face Inference Endpoints on Anchor Terminal, B, 64.5/100
    On a light page
    On a dark page
    <a href="https://www.anchorterminal.com/tools/hugging-face-inference-endpoints"><img src="https://www.anchorterminal.com/badges/hugging-face-inference-endpoints.svg" alt="Hugging Face Inference Endpoints on Anchor Terminal" height="20"></a>
    [![Hugging Face Inference Endpoints on Anchor Terminal](https://www.anchorterminal.com/badges/hugging-face-inference-endpoints.svg)](https://www.anchorterminal.com/tools/hugging-face-inference-endpoints)

    It counts on the README of github.com/huggingface/hf-endpoints-documentation.

  2. Tell us where it is

    We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.

Agents send the same to POST /api/v1/verify as {"slug": "hugging-face-inference-endpoints", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.