# Hugging Face Inference Endpoints > Managed Hugging Face service that deploys a Hub model as a dedicated, autoscaling HTTPS endpoint on AWS, Azure or Google Cloud, using vLLM, TGI, SGLang, llama.cpp, TEI or a custom container. Managed by REST API, Python client, CLI or MCP. - Canonical: https://www.anchorterminal.com/tools/hugging-face-inference-endpoints - Markdown: https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.md (~8,600 tokens) - Slim: https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.min.md (~1,980 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 ## Overview **Grade B · 64.5/100 · rank #314 of 842 · #3 in GPU & serverless compute · not agent-ready · confidence medium** ## Assessment OAuth scopes separate reading endpoints from writing them, both OpenAPI documents are public, and the unauthenticated `/v2/provider` route lists every instance with its hourly price. An account needs a payment method and credits before the first deployment, no rate limits or SLA were found for the management API, and the docs price table disagrees with the live list in places. ## Facts | Field | Value | | --- | --- | | Vendor | Hugging Face, Inc. (https://huggingface.co) | | Kind | HTTP API | | Category | GPU & serverless compute (https://www.anchorterminal.com/categories/gpu-compute) | | Transport | HTTP | | Endpoint | `https://api.endpoints.huggingface.cloud` | | Auth | OAuth or key · Hugging Face access token sent as `Authorization: Bearer $HF_TOKEN` to the management API and to each endpoint. Tokens are created in the account settings in a browser and can be fine-grained, read or write. The MCP server uses OAuth through huggingface.co (authorisation code with PKCE, device code, dynamic client registration) with `read-endpoints` and `write-endpoints` scopes. `GET /v2/provider` and the catalogue list need no token. Access is self-serve, with quota requests for larger instances. | | Pricing | Pay per use ($0.033 / vCPU-hr) · Usage priced by instance hour, billed per minute while a replica is initialising or running. GPUs run from $0.50 an hour (T4) to $10 (H100 on GCP), CPUs from $0.033. No free tier or trial was found. The docs require a payment method and credits before deploying, and the pricing page says an active subscription. Paused endpoints and endpoints at zero replicas aren't billed for compute (https://huggingface.co/docs/inference-endpoints/support/pricing). | | x402 | No · No x402, MPP or L402 in the Inference Endpoints docs, the two OpenAPI documents or the pricing page (checked 2026-10-08). | | Licence | Proprietary service under the Hugging Face Terms of Service. The `huggingface_hub` Python client and CLI are Apache-2.0 | | Tools exposed | 19 | | Packages | pypi: `huggingface_hub` | | Source | https://github.com/huggingface/hf-endpoints-documentation | | Docs | https://huggingface.co/docs/inference-endpoints/index | | llms.txt | https://huggingface.co/docs/inference-endpoints/llms.txt | | Last release | 2026-10-08 | | PyPI downloads / week | 60,014,944 | | Surfaces | Management REST API at https://api.endpoints.huggingface.cloud (paths under `/v2` and `/v3`), catalogue API at https://endpoints.huggingface.co/api/v1, MCP server at https://endpoints.huggingface.co/mcp, the `huggingface_hub` Python client and the `hf endpoints` CLI | | Engines | vLLM, Text Generation Inference (in maintenance mode since 11 December 2025), SGLang, llama.cpp, Text Embeddings Inference, the Inference Toolkit, or a custom container listening on port 80 | | GPUs | T4, L4, A10G, L40S, A100, RTX PRO 6000 Blackwell, H100 (GCP), H200 (GCP), AWS Inferentia2, and CPU instances, in sizes x1 to x8 | | Clouds and regions | AWS us-east-1, us-east-2 and eu-west-1, Azure eastus, GCP us-east4 and us-south1, per the live provider list. AWS us-west-2 is listed as not available | | Scale to zero | Optional, with `minReplica` 0. The docs give a default idle period of 1 hour and the OpenAPI document says 15 minutes. `POST /v2/endpoint/{namespace}/{name}/scale-to-zero` forces it | | Cold start | No figure published. The docs say initialising usually takes 3 to 5 minutes and that scaling up can take a few minutes. The proxy returns 503 meanwhile unless the request carries `X-Scale-Up-Timeout` | | Autoscaling | By hardware utilisation (default threshold 80 per cent) or pending requests (default 1.5 a replica over 20 seconds). Scale-up is evaluated every minute, scale-down every 2 minutes with a 300-second stabilisation | | Billing basis | Hourly rate billed per minute for each replica while initialising or running. Paused endpoints stop billing. Prepaid credits, with optional automatic recharge | | Endpoint access | Private (default, Hugging Face token of the owner or organisation members), authenticated (any Hugging Face token) or public. AWS PrivateLink is available on AWS | | MCP tools | 19, including `list_endpoints`, `create_endpoint`, `update_endpoint`, `pause_endpoint`, `scale_endpoint_to_zero`, `delete_endpoint`, `call_endpoint`, `get_recommended_config`, `get_endpoint_logs`, `get_endpoint_metric` and `get_audit_logs` | | Data retention | The docs say request payloads and tokens passed to an endpoint aren't stored and logs are kept for 30 days. A GDPR data processing agreement comes with an Enterprise subscription | | Capabilities | compute.gpu, compute.endpoints, compute.containers | | Tags | hosted, usage-priced, python, cli, mcp, oauth, openapi, llms-txt, status-page, soc2, enterprise | | JSON | https://www.anchorterminal.com/api/v1/tools/hugging-face-inference-endpoints.json | ## Score breakdown (methodology v0.4, October 2026 research run) Assessed 2026-10-08 from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/#checklist). Confidence: medium. Performance and Task success pending (no score, not in the total); the total is Σ(score × weight) ÷ 80 over the 7 assessed categories. "This run" is each category's share of the 100 points. | Category | Weight | This run | Score (0–100) | Points | | --- | --- | --- | --- | --- | | Reliability | 16% | 20 | 63 | 12.6 | | Performance | 10% | pending | pending | n/a | | Schema & documentation | 13% | 16.2 | 73 | 11.9 | | Agent ergonomics | 13% | 16.2 | 62 | 10.1 | | Security & auth | 14% | 17.5 | 83 | 14.5 | | Payments & pricing | 10% | 12.5 | 20 | 2.5 | | Task success | 10% | pending | pending | n/a | | Maintenance & community | 7% | 8.8 | 80 | 7.0 | | Transparency & trust (editorial 68, provenance 67) | 7% | 8.8 | 68 | 6.0 | | Negative events | up to −15 | up to −15 | none recorded | 0 | | **Total** | | | | **64.5 → B** | ### Why each score - Reliability 63: Better Stack status page at status.huggingface.co with separate Inference Endpoints UI and API components and 90 days of history (20). The API component shows 100 per cent uptime and three scheduled maintenance windows (22 and 30 July, 7 October 2026). The UI was down for 1 hour 37 minutes on 16 July 2026 during a Hub outage, which we count as minor for the API surface (20). No rate limits were found for the management API. The Hub publishes limits per 5-minute window (1,000 API requests for a free user), without saying they cover api.endpoints.huggingface.cloud (5). The docs explain the 503 returned while a replica starts and the `X-Scale-Up-Timeout` header that holds the request, and advise 2 replicas for availability. No backoff guidance or idempotency keys for writes (8). No SLA found in the docs or on the pricing page (0). The service is generally available (10). - Performance: Pending. Latency is measured per call by our probes, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until the first probe window closes. - Schema & documentation 73: Public OpenAPI 3.1 documents for the management API (40 paths, 46 operations) and the catalogue API (3 operations) (25). llms.txt with a Markdown twin of every docs page (10). Every operation has a summary tagged READ, WRITE, PUBLIC or PRO, but only 5 of 46 have a description. The MCP tool table states each tool's purpose and says to call `get_recommended_config` before `create_endpoint` (11). 127 typed schemas with required fields and enums for state, type and accelerator. List filters are comma-separated strings (12). Parameters carry examples. The management document lists only 200 responses, plus 501 on three log routes, while the catalogue document lists 400, 401, 404, 409 and 500 (7). Paths are versioned `/v2` and `/v3` and the document is version 2.0.0. Inference Endpoints has no changelog of its own. The Hub changelog and the docs repository history are the dated record (8). - Agent ergonomics 62: List endpoints takes `limit` (default 20) and `cursor`, logs take `limit`, `tail` and `line_max_length`, and the MCP server has 19 tools (18). Cursor pagination, filters by state, type, task and tags, sorting, and a time window on logs (18). Errors are documented for the catalogue API and in prose for pause, resume and scale to zero. The management document has no error schema (8). No idempotency keys. Endpoints are addressed by name, updates are a PUT, and the MCP delete tool previews before a confirmed second call. MCP annotations weren't read (8). One-call deploy from the catalogue with tuned defaults, a Python client and the `hf endpoints` CLI. A second-language SDK for endpoint management wasn't found (10). - Security & auth 83: OAuth through huggingface.co with `read-endpoints` and `write-endpoints` scopes, PKCE and dynamic client registration for the MCP server, and revocable fine-grained access tokens for the API. No documented option to send a token in a URL (30). Read and write scopes are separate, endpoints are private by default, organisations on Team and Enterprise plans can require approval of fine-grained tokens, and the MCP delete tool needs `confirm: true`. Other write calls have no confirmation step (17). The service returns the output and logs of the owner's own model, not third-party content (10). An audit log route and MCP tool on Pro, Team or Enterprise plans, plus logs and metrics for every endpoint (12). security.txt valid to 2030 with security@huggingface.co, SOC2 Type 2 stated for the Hub and Inference Endpoints, malware and pickle scanning of repositories. No bug bounty was found (14). - Payments & pricing 20: No machine payment protocol (0). Hourly prices for every instance published without a login and returned by the unauthenticated `/v2/provider` route (20). No free tier or trial. The docs require a payment method and credits (0). Signup, billing and token creation are browser steps (0). - Task success: Pending. Task success needs the category task suites run through each tool, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until then. A data provider's data-quality score is published on its listing now and becomes half of this category when it's scored. - Maintenance & community 80: `huggingface_hub` 2.2.0 was published on PyPI on 8 October 2026 and the docs repository changed on 7 October (30). Releases 1.31.0 to 2.2.0 in the 30 days before, about one a week, and 8 docs commits since 10 July 2026 (20). The GitHub API refused our request, so issue response times went unread. A public changelog for the Hub, a forum and a quota contact form exist (8 of 15). Current official Python client and CLI (15). The client supports Python 3.10 and later and ships weekly. Its CI wasn't read (7). - Transparency & trust 68: Closed service under clear terms that name Inference Endpoints, with an Apache-2.0 client (20). The security page says payloads and tokens aren't stored and logs are kept 30 days. The privacy policy dates from 28 March 2023 and gives retention as long as necessary, and a data processing agreement requires an Enterprise subscription (20). Instance deprecations are dated to the month in the pricing table and flagged in the provider API, and the TGI maintenance notice is dated. No deprecation policy for the API was found (10). The privacy policy lists 11 subprocessors with countries and the docs and API list deployment regions (18). Fix list for a coding agent, everything this grade says the listing lacks, the biggest gain first (19 items): https://www.anchorterminal.com/fixes/hugging-face-inference-endpoints.md (JSON https://www.anchorterminal.com/fixes/hugging-face-inference-endpoints.json) ### What we couldn't check - unchecked: GitHub stars, issues and CI for huggingface_hub, because api.github.com answered 403 - unchecked: the Supplemental Terms PDF beyond its first page, which is all our reader extracted - unchecked: the MCP tool definitions and annotations, because the server needs an OAuth login - unchecked: the permissions a fine-grained token can carry for Inference Endpoints, which are shown only in the logged-in token settings - unchecked: security incidents and advisories in the last 12 months, with web search unavailable - unchecked: the registration date of huggingface.co, because RDAP returned 404 - No SLA, rate limit or idempotency documentation was found for the management API. An Enterprise contract may carry an SLA that isn't published. - The docs price table and the live `/v2/provider` list disagree on Inferentia2 x1 ($0.75 against $1.95), AWS H200 availability and the RTX PRO 6000 Blackwell, and on the default scale-to-zero period (1 hour against 15 minutes). No deduction was taken. - The lead was right on the product and interfaces. It omitted the MCP server and the catalogue API. Scale to zero and prices are now read. ### Sources - docs index (llms.txt): (seen 2026-10-08) - pricing: (seen 2026-10-08) - FAQ: (seen 2026-10-08) - autoscaling guide: (seen 2026-10-08) - configuration guide: (seen 2026-10-08) - security and compliance: (seen 2026-10-08) - MCP server guide: (seen 2026-10-08) - API reference page: (seen 2026-10-08) - management OpenAPI document: (seen 2026-10-08) - live provider list: (seen 2026-10-08) - catalogue OpenAPI document: (seen 2026-10-08) - MCP OAuth resource metadata: (seen 2026-10-08) - OAuth authorisation server metadata: (seen 2026-10-08) - status page: (seen 2026-10-08) - access tokens: (seen 2026-10-08) - Hub rate limits: (seen 2026-10-08) - Python client guide: (seen 2026-10-08) - huggingface_hub on PyPI: (seen 2026-10-08) - PyPI downloads: (seen 2026-10-08) - docs repository history: (seen 2026-10-08) - terms of service: (seen 2026-10-08) - privacy policy: (seen 2026-10-08) - security.txt: (seen 2026-10-08) - platform pricing: (seen 2026-10-08) - Hub changelog: (seen 2026-10-08) ## Who's behind it (provenance 67/100, checked 2026-10-08) | Check | Finding | Points | | --- | --- | --- | | Legal entity named | Hugging Face, Inc. | 20/20 | | Domain age | huggingface.co, no registry record we could read | 0/15 | | Endpoint on the vendor's domain | api.endpoints.huggingface.cloud is not on huggingface.co | 0/15 | | Terms of service | read, states 6 of the 7 things a reader expects, and has 1 clause that costs points | 7.1/10 | | Privacy policy | read, states 8 of the 8 things a reader expects | 10/10 | | Status page | status.huggingface.co | 10/10 | | Changelog | published | 10/10 | | security.txt | valid | 10/10 | The Terms of Service (effective 15 September 2022) name Hugging Face, Inc., a Delaware corporation, list Inference Endpoints among the services they cover and are governed by New York law. They link Supplemental Terms (effective 28 April 2025) as a PDF, of which our reader extracted only the first page. The privacy policy (effective 28 March 2023) names Hugging Face, Inc. and its EU establishment Hugging Face, SAS, 9 rue des Colonnes, 75002 Paris, and lists 11 subprocessors with countries. The Inference Endpoints security page points to it. The management API answers at api.endpoints.huggingface.cloud and deployed endpoints at subdomains of endpoints.huggingface.cloud, a second domain of the vendor's. The catalogue API and the MCP server are on endpoints.huggingface.co. huggingface.co/.well-known/security.txt gives security@huggingface.co and expires on 1 July 2030. endpoints.huggingface.co/.well-known/security.txt returns 404. status.huggingface.co is a Better Stack page with separate components for the Inference Endpoints UI and API. The changelog at huggingface.co/changelog covers the whole Hub. Inference Endpoints has no changelog of its own. Dated changes are in the docs repository's commit history. rdap.org returned 404 for huggingface.co, so the registration date is unrecorded. ### Terms and privacy, as read A reading by a fixed set of rules, each answered with the vendor's own sentence. Not legal advice. **Terms of service** (https://huggingface.co/terms-of-service), read 2026-10-08, dated 2022-09-15, states 6 of the 7 things a reader expects. - To know. Says the terms or the service can change without notice (costs points). "We may at any time modify, suspend, or discontinue, temporarily or permanently, the Services (or any part thereof) with or without notice." - To know. Says access can be ended without notice or for any reason. "We may do the same, and we reserve the right to suspend or terminate your access to the Services anytime with or without cause, and at our own discretion, with or without notice." - To know. Has not been updated for three years or more. "🗓 Effective Date: September 15, 2022" - Gives the date it was last updated. Last updated 2022-09-15. - Names the governing law or courts. The law of the State of New York. - States a limit on its liability. Capped at the fees paid in the 12 months before the claim or $50. - Says how changes to the terms are announced. Changes are posted, with no other notice named. - Not found in the text. Refers to a service level or uptime commitment. - Also in the text (2026-10-08). Each party's total liability is capped at the amount paid in the 12 months before the last claim, or 50 US dollars where the claim relates to a free service. "Either Party’s (and each Related Party’s) aggregate liability to the other Party or any third party in any circumstance will not exceed the amount that you paid us during the 12-month period immediately preceding the last claim (or $50 if relating to a free service)." - Also in the text (2026-10-08). Setting a repository public grants every user a perpetual, irrevocable licence to use, reproduce, distribute and make derivative works of its content through the services. "If you decide to set your Repository public, you grant each User a perpetual, irrevocable, worldwide, royalty-free, non-exclusive license to use, display, publish, reproduce, distribute, and make derivative works of your Content through our Services and functionalities;" - Also in the text (2026-10-08). After an account is cancelled, the vendor says it will use commercially reasonable efforts to delete the account's information and repository content within 90 days. "Upon cancellation of your Account, we will use commercially reasonable efforts to delete your information and Content of your own Repositories, whether public or private, within 90 days." **Privacy policy** (https://huggingface.co/privacy), read 2026-10-08, dated 2023-03-28, states 8 of the 8 things a reader expects. - To know. Has not been updated for three years or more. "🗓 Effective Date: March 28, 2023" - Gives the date it was last updated. Last updated 2023-03-28. - Says how long data is kept. For as long as needed, with no period named. - Says whether personal data is sold or shared for advertising. Says it does not sell personal data. - Gives a privacy contact. privacy@huggingface.co. - Also in the text (2026-10-08). Aggregated information that does not identify a user may be disclosed to advertisers and partners, with or without payment, for purposes that include targeting advertisements. "The Company may disclose Anonymous Information (with or without compensation) to third parties, including advertisers and partners, for purposes including, but not limited to, targeting advertisements." - Also in the text (2026-10-08). The company may access information a user keeps private without consent, for legitimate interests such as maintaining security or meeting legal and regulatory obligations. "The Company also reserves the right to access this information with your consent, or without your consent only for the purposes of pursuing legitimate interests such as maintaining security on its Services or complying with any legal or regulatory obligations." ## Live (updated 2026-10-09 11:46 UTC) - Right now: up, HTTP 200, 257 ms, checked 2026-10-09 11:46 UTC (get on `https://api.endpoints.huggingface.cloud`) - Uptime 24h 100.0% (44 probes) · 30 days 100.0% (44 probes) · p50 255 ms · p95 300 ms - Vendor status page: unknown, no machine-readable status found - Always current: https://www.anchorterminal.com/api/v1/live/hugging-face-inference-endpoints.json ## Probe metrics Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. Live uptime, where we poll the endpoint, is under Live and doesn't change the score. ## Prices | Item | Price | Unit | Note | | --- | --- | --- | --- | | NVIDIA T4 16 GB x1 (AWS, GCP) | $0.50 | per GPU-hour | Billed per minute while initialising or running | | NVIDIA L4 24 GB x1 (AWS) | $0.80 | per GPU-hour | $0.70 on GCP us-east4 | | NVIDIA A10G 24 GB x1 (AWS) | $1 | per GPU-hour | us-east-1 and eu-west-1 | | NVIDIA L40S 48 GB x1 (AWS) | $1.80 | per GPU-hour | us-east-1 | | NVIDIA A100 80 GB x1 (AWS) | $2.50 | per GPU-hour | $3.60 on GCP us-east4 | | NVIDIA RTX PRO 6000 Blackwell 96 GB x1 (AWS) | $2.75 | per GPU-hour | us-east-2, in the live provider list and absent from the docs table | | NVIDIA H200 141 GB x1 (GCP) | $5 | per GPU-hour | us-south1. The AWS H200 in us-west-2 is marked deprecated in the API | | NVIDIA H100 80 GB x1 (GCP) | $10 | per GPU-hour | us-east4. The AWS H100 at $4.50 is deprecated from December 2025 | | Intel Sapphire Rapids x1, 1 vCPU and 2 GB (AWS) | $0.033 | per vCPU-hour | $0.050 on GCP and $0.060 on Azure Intel Xeon | Across all listings: https://www.anchorterminal.com/prices/index.md ## Strengths - Public OpenAPI 3.1 documents for the management API (46 operations) and the catalogue API (3), plus llms.txt and a Markdown twin of every docs page - The MCP server at endpoints.huggingface.co/mcp uses OAuth with `read-endpoints` and `write-endpoints` scopes, PKCE and dynamic client registration - `GET /v2/provider` needs no token and returns each instance type by cloud and region with status and price per hour - The MCP `delete_endpoint` tool returns a preview and deletes only on a second call with `confirm: true` - The Inference Endpoints API component on status.huggingface.co shows 100 per cent uptime over the 90 days to 8 October 2026 ## Weaknesses - No free tier. The docs require a payment method and credits, and replicas are billed while initialising as well as running - No rate limits, SLA or idempotency keys were found for the management API, and its OpenAPI document lists only 200 responses on 43 of 46 operations - The docs price table and the live provider list disagree. Inferentia2 x1 is $0.75 in the docs and $1.95 in the API, and AWS H200 is listed in the docs and marked deprecated in the API - A start from zero replicas takes minutes by the docs' own account, and the proxy answers 503 until a replica is ready - The Hub outage of 16 July 2026 took the Inference Endpoints UI down for 1 hour 37 minutes ## Before you call it (notes for agents) 1. Call `GET https://api.endpoints.huggingface.cloud/v2/provider` first and pick an instance whose `status` is `available`. The docs table lists types the API marks deprecated or not available 2. Send `X-Scale-Up-Timeout: 600` on requests to an endpoint that scales to zero, or handle 503 while the first replica starts 3. Set `scaleToZeroTimeout` yourself. The docs give a default of 1 hour and the OpenAPI document says 15 minutes 4. Pause or delete an endpoint when the job is done. Billing covers every minute a replica is initialising or running 5. Give the agent a fine-grained token or the `read-endpoints` scope unless it must deploy. Endpoints are private by default and take the same Hugging Face token as a bearer ## Connect Install: ```bash pip install huggingface_hub ``` First request: ```bash curl "https://api.endpoints.huggingface.cloud/v2/endpoint/$NAMESPACE" \ -H "Authorization: Bearer $HF_TOKEN" ``` Through letme (picks today, calling later): https://letme.dev/hugging-face-inference-endpoints. letme answers with the pick and how to call it direct; calling through letme (one key, the vendor's own price) comes later. How it works: https://www.anchorterminal.com/letme/index.md ## Similar tools Ranked by shared capabilities, then score. Same-category tools with no shared capability key are listed last. | Tool | Grade | Score | Rank | Shared capabilities | x402 | Markdown | | --- | --- | --- | --- | --- | --- | --- | | Nebius AI Cloud | B | 67.2 | 240 | compute.gpu, compute.endpoints, compute.containers | no | https://www.anchorterminal.com/tools/nebius-ai-cloud.md | | Baseten | B | 66.5 | 264 | compute.gpu, compute.endpoints, compute.containers | no | https://www.anchorterminal.com/tools/baseten.md | | Modal | B | 63.6 | 346 | compute.gpu, compute.endpoints, compute.containers | no | https://www.anchorterminal.com/tools/modal.md | | Replicate Deployments | B | 63.6 | 347 | compute.gpu, compute.endpoints, compute.containers | no | https://www.anchorterminal.com/tools/replicate-deploy.md | | Vast.ai | B | 62.6 | 384 | compute.gpu, compute.containers, compute.endpoints | no | https://www.anchorterminal.com/tools/vast-ai.md | | Verda | B | 62.3 | 396 | compute.gpu, compute.endpoints, compute.containers | no | https://www.anchorterminal.com/tools/verda.md | ## Panel reviews (0) Reviewed by the Anchor panel (https://www.anchorterminal.com/reviewers/index.md): . Desk reviews, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. For a desk review, the outcome says whether the reviewer's questions could be answered from public material: success, partial or failure. How reviews work: https://www.anchorterminal.com/reviews/how-it-works.md ## Notable - The management OpenAPI 3.1 document (title HF Inference Endpoints API, version 2.0.0) has 40 paths and 46 operations, each summary tagged [READ], [WRITE], [PUBLIC] or [PRO] (source: ) - A separate catalogue API under `/api/v1` on endpoints.huggingface.co lists catalogue models without a token and deploys a model or a named recipe in one POST, with 400, 401, 404, 409 and 500 responses documented (source: ) - The MCP server documentation was added on 19 August 2026 and lists 19 tools. `get_audit_logs` needs a Pro or Enterprise plan (source: ) - The MCP resource metadata names huggingface.co as the authorisation server and the scopes openid, profile, email, read-repos, read-billing, inference-api, read-endpoints and write-endpoints (source: ) - The pricing page says accounts need an active subscription and credits added, and that replicas are charged while initialising and running, by the minute (source: ) - The pricing table marks AWS H100 and B200 as deprecated from December 2025 and AWS Ice Lake CPUs from July 2025. GCP H200 was added to the table on 7 October 2026 (source: ) - The security page says the Hub and Inference Endpoints are SOC2 Type 2 certified, payloads and tokens aren't stored, and logs are kept for 30 days (source: ) - The FAQ advises at least 2 replicas for an endpoint that must stay available, after intermittent 503 errors on running endpoints (source: ) - huggingface.co/pricing describes Inference Endpoints as having no cold starts, while the autoscaling guide describes a cold start after scale to zero (source: ) ## Compare - [Baseten vs Hugging Face Inference Endpoints](https://www.anchorterminal.com/compare/baseten-vs-hugging-face-inference-endpoints.md): B 66.5 vs B 64.5 - [Beam vs Hugging Face Inference Endpoints](https://www.anchorterminal.com/compare/beam-vs-hugging-face-inference-endpoints.md): C 55.5 vs B 64.5 - [Cerebrium vs Hugging Face Inference Endpoints](https://www.anchorterminal.com/compare/cerebrium-vs-hugging-face-inference-endpoints.md): C 55.3 vs B 64.5 - [CoreWeave vs Hugging Face Inference Endpoints](https://www.anchorterminal.com/compare/coreweave-vs-hugging-face-inference-endpoints.md): C 61.5 vs B 64.5 - [Hugging Face Inference Endpoints vs Hyperbolic](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-hyperbolic.md): B 64.5 vs D 48 - [Hugging Face Inference Endpoints vs Koyeb](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-koyeb.md): B 64.5 vs D 46.5 - [Hugging Face Inference Endpoints vs Lambda Cloud](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-lambda.md): B 64.5 vs D 50 - [Hugging Face Inference Endpoints vs Modal](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-modal.md): B 64.5 vs B 63.6 - [Hugging Face Inference Endpoints vs Nebius AI Cloud](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-nebius-ai-cloud.md): B 64.5 vs B 67.2 - [Hugging Face Inference Endpoints vs Northflank](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-northflank.md): B 64.5 vs C 61.3 - [Hugging Face Inference Endpoints vs Replicate Deployments](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-replicate-deploy.md): B 64.5 vs B 63.6 - [Hugging Face Inference Endpoints vs Runpod](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-runpod.md): B 64.5 vs D 53.5 - [Hugging Face Inference Endpoints vs Thunder Compute](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-thunder-compute.md): B 64.5 vs C 56.1 - [Hugging Face Inference Endpoints vs Vast.ai](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-vast-ai.md): B 64.5 vs B 62.6 - [Hugging Face Inference Endpoints vs Verda](https://www.anchorterminal.com/compare/hugging-face-inference-endpoints-vs-verda.md): B 64.5 vs B 62.3 ## Verify this listing For the vendor. The badge or a plain link to this page verifies the listing, from the README of github.com/huggingface/hf-endpoints-documentation. It shows the listing is the vendor's and that the vendor knows it's here, and it never changes a grade, rank or review. The vendor sends the page's address to `POST https://www.anchorterminal.com/api/v1/verify` as `{"slug": "hugging-face-inference-endpoints", "url": "…"}`, or calls the `verify_listing` tool at https://www.anchorterminal.com/mcp. We fetch the page once, then again every week; two failed checks in a row and the verification lapses, and a later pass restores it. What we check: https://www.anchorterminal.com/builders/index.md#verify HTML badge: ```html Hugging Face Inference Endpoints on Anchor Terminal ``` Markdown badge, for a README: ```markdown [![Hugging Face Inference Endpoints on Anchor Terminal](https://www.anchorterminal.com/badges/hugging-face-inference-endpoints.svg)](https://www.anchorterminal.com/tools/hugging-face-inference-endpoints) ``` Plain link: ```html Hugging Face Inference Endpoints on Anchor Terminal ``` ## Share this listing For the vendor. Sharing assets for social media, two PNGs of 1200 × 630 that say Hugging Face Inference Endpoints is listed on Anchor Terminal, with the vendor's logo and this page's address and no grade or score. - Dark: https://www.anchorterminal.com/assets/share/hugging-face-inference-endpoints-dark.png - Light: https://www.anchorterminal.com/assets/share/hugging-face-inference-endpoints-light.png