# Hugging Face Inference Endpoints (slim) > Managed Hugging Face service that deploys a Hub model as a dedicated, autoscaling HTTPS endpoint on AWS, Azure or Google Cloud, using vLLM, TGI, SGLang, llama.cpp, TEI or a custom container. Managed by REST API, Python client, CLI or MCP. - Full: https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.md (~8,600 tokens) · this version ~1,980 tokens · JSON https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.json · canonical https://www.anchorterminal.com/tools/hugging-face-inference-endpoints - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 64.5/100 · rank #314 of 842 · #3 in GPU & serverless compute · not agent-ready · confidence medium** Assessment: OAuth scopes separate reading endpoints from writing them, both OpenAPI documents are public, and the unauthenticated `/v2/provider` route lists every instance with its hourly price. An account needs a payment method and credits before the first deployment, no rate limits or SLA were found for the management API, and the docs price table disagrees with the live list in places. ## Facts - Kind: HTTP API · vendor: Hugging Face, Inc. · category: GPU & serverless compute · legal entity: Hugging Face, Inc. · provenance 67/100 - Endpoint: `https://api.endpoints.huggingface.cloud` (HTTP) - Auth: OAuth or key · pricing: Pay per use · x402: no · licence: Proprietary service under the Hugging Face Terms of Service. The `huggingface_hub` Python client and CLI are Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Surfaces: Management REST API at https://api.endpoints.huggingface.cloud (paths under `/v2` and `/v3`), catalogue API at https://endpoints.huggingface.co/api/v1, MCP server at https://endpoints.huggingface.co/mcp, the `huggingface_hub` Python client and the `hf endpoints` CLI - Engines: vLLM, Text Generation Inference (in maintenance mode since 11 December 2025), SGLang, llama.cpp, Text Embeddings Inference, the Inference Toolkit, or a custom container listening on port 80 - GPUs: T4, L4, A10G, L40S, A100, RTX PRO 6000 Blackwell, H100 (GCP), H200 (GCP), AWS Inferentia2, and CPU instances, in sizes x1 to x8 - Clouds and regions: AWS us-east-1, us-east-2 and eu-west-1, Azure eastus, GCP us-east4 and us-south1, per the live provider list. AWS us-west-2 is listed as not available - Scale to zero: Optional, with `minReplica` 0. The docs give a default idle period of 1 hour and the OpenAPI document says 15 minutes. `POST /v2/endpoint/{namespace}/{name}/scale-to-zero` forces it - Cold start: No figure published. The docs say initialising usually takes 3 to 5 minutes and that scaling up can take a few minutes. The proxy returns 503 meanwhile unless the request carries `X-Scale-Up-Timeout` - Autoscaling: By hardware utilisation (default threshold 80 per cent) or pending requests (default 1.5 a replica over 20 seconds). Scale-up is evaluated every minute, scale-down every 2 minutes with a 300-second stabilisation - Billing basis: Hourly rate billed per minute for each replica while initialising or running. Paused endpoints stop billing. Prepaid credits, with optional automatic recharge - Endpoint access: Private (default, Hugging Face token of the owner or organisation members), authenticated (any Hugging Face token) or public. AWS PrivateLink is available on AWS - MCP tools: 19, including `list_endpoints`, `create_endpoint`, `update_endpoint`, `pause_endpoint`, `scale_endpoint_to_zero`, `delete_endpoint`, `call_endpoint`, `get_recommended_config`, `get_endpoint_logs`, `get_endpoint_metric` and `get_audit_logs` - Data retention: The docs say request payloads and tokens passed to an endpoint aren't stored and logs are kept for 30 days. A GDPR data processing agreement comes with an Enterprise subscription - Prices: NVIDIA T4 16 GB x1 (AWS, GCP) $0.50 per GPU-hour; NVIDIA L4 24 GB x1 (AWS) $0.80 per GPU-hour; NVIDIA A10G 24 GB x1 (AWS) $1 per GPU-hour; NVIDIA L40S 48 GB x1 (AWS) $1.80 per GPU-hour; NVIDIA A100 80 GB x1 (AWS) $2.50 per GPU-hour; NVIDIA RTX PRO 6000 Blackwell 96 GB x1 (AWS) $2.75 per GPU-hour; NVIDIA H200 141 GB x1 (GCP) $5 per GPU-hour; NVIDIA H100 80 GB x1 (GCP) $10 per GPU-hour; Intel Sapphire Rapids x1, 1 vCPU and 2 GB (AWS) $0.033 per vCPU-hour - Scores: Reliability 63, Performance pending, Schema & documentation 73, Agent ergonomics 62, Security & auth 83, Payments & pricing 20, Task success pending, Maintenance & community 80, Transparency & trust 68 · total over the 7 assessed categories - Why: Reliability, Better Stack status page at status.huggingface.co with separate Inference Endpoints UI and API components and 90 days of history (20). · Schema & documentation, Public OpenAPI 3.1 documents for the management API (40 paths, 46 operations) and the catalogue API (3 operations) (25). · Agent ergonomics, List endpoints takes `limit` (default 20) and `cursor`, logs take `limit`, `tail` and `line_max_length`, and the MCP server has 19 tools (18… · Security & auth, OAuth through huggingface.co with `read-endpoints` and `write-endpoints` scopes, PKCE and dynamic client registration for the MCP server, an… · Payments & pricing, No machine payment protocol (0). · Maintenance & community, `huggingface_hub` 2.2.0 was published on PyPI on 8 October 2026 and the docs repository changed on 7 October (30). · Transparency & trust, Closed service under clear terms that name Inference Endpoints, with an Apache-2.0 client (20). - Sources: 25, open questions: 9, both in the full twin - Capabilities: compute.gpu, compute.endpoints, compute.containers - JSON: https://www.anchorterminal.com/api/v1/tools/hugging-face-inference-endpoints.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/hugging-face-inference-endpoints.svg` or a link to https://www.anchorterminal.com/tools/hugging-face-inference-endpoints from the README of github.com/huggingface/hf-endpoints-documentation, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `GET https://api.endpoints.huggingface.cloud/v2/provider` first and pick an instance whose `status` is `available`. The docs table lists types the API marks deprecated or not available 2. Send `X-Scale-Up-Timeout: 600` on requests to an endpoint that scales to zero, or handle 503 while the first replica starts 3. Set `scaleToZeroTimeout` yourself. The docs give a default of 1 hour and the OpenAPI document says 15 minutes 4. Pause or delete an endpoint when the job is done. Billing covers every minute a replica is initialising or running 5. Give the agent a fine-grained token or the `read-endpoints` scope unless it must deploy. Endpoints are private by default and take the same Hugging Face token as a bearer ## Connect ```bash pip install huggingface_hub ``` ```bash curl "https://api.endpoints.huggingface.cloud/v2/endpoint/$NAMESPACE" \ -H "Authorization: Bearer $HF_TOKEN" ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/hugging-face-inference-endpoints ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Nebius AI Cloud | B | 67.2 | compute.gpu, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/nebius-ai-cloud.min.md | | Baseten | B | 66.5 | compute.gpu, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/baseten.min.md | | Modal | B | 63.6 | compute.gpu, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/modal.min.md | | Replicate Deployments | B | 63.6 | compute.gpu, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/replicate-deploy.min.md | | Vast.ai | B | 62.6 | compute.gpu, compute.containers, compute.endpoints | https://www.anchorterminal.com/tools/vast-ai.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)