> For the complete documentation index, see [llms.txt](https://docs.obitmc.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.obitmc.com/inference-api/where-your-prompt-runs.md).

# Where your prompt runs

Where an Obit request runs: which providers serve the public pool, what a worker machine can and cannot see, how requests are encrypted, and dedicated endpoints.

Where your prompt goes and who controls that machine. For what we keep, the binding statement is the [Privacy Policy (Inference)](/privacy-policy-inference.md).

*Last verified: 2026-09-22*

## The short version

Your request reaches an edge relay over HTTPS, which hands it to one worker over an authenticated, encrypted connection. The worker generates the answer in memory, streams it back the same way, and is never reachable from the internet.

## Which machines serve the public pool

**Today the public pool is served primarily from** [**Vast.ai**](https://vast.ai), and we scale across our providers as demand requires. Every successful response carries an `X-Obit-Provider` header naming the provider that served it. If you need a particular provider for a model, write to <jboesch@obitmc.com> with your use case and volume and we will set it up.

### Pinning a request to a provider

Pin a request with a `<model>:<selector>` suffix or a `provider: {"only": […]}` object. A pin is **strict**: if nothing in the pinned set can serve the request you get a `503 no_provider_match`, never a request quietly served somewhere else. Read the live selectors from `GET /v1/models` (`providers` and `provider_groups`). Pin only for a hard requirement, and see [Models and pricing](/inference-api/models-and-pricing.md) for the rules and the cost of pinning.

## What the worker machine can and cannot see

* **It can see the request it is serving.** Prompt and completion sit in that machine's memory while the request runs.
* **A prefix of your prompt may stay in its cache while it is running**, which is what makes a repeated prefix cheaper. It is reported as `usage.prompt_tokens_details.cached_tokens`, is held in memory only, is never written to disk, and dies with the worker.
* **It is not reachable from the internet.** Workers accept no inbound connections and dial out to the relay themselves.
* **Workers run on hosted infrastructure.** As with any cloud provider, the host operates the physical machine while we control the software and your data. Where a stricter boundary is required, [dedicated endpoints](#dedicated-and-isolated-endpoints) are available.

For what we keep, see the [Privacy Policy (Inference)](/privacy-policy-inference.md).

## Encryption

* **You to the edge**: HTTPS/TLS to `relay.obitmc.com`, terminated on Cloudflare.
* **The edge to a worker**: an authenticated WebSocket Secure (`wss://`) session the worker opens, validated before anything is routed to it.
* **Between our own services**: TLS inside a private cluster, each hop authenticated.

No hop is plaintext.

## Dedicated and isolated endpoints

If a shared pool is not acceptable, we run **dedicated endpoints**: a pool that serves only you, on hardware chosen for your requirements, at your own endpoint. Email <jboesch@obitmc.com> with what you run and the constraint you need to satisfy.

## Status and incidents

[**status.obitmc.com**](https://status.obitmc.com) shows per-model availability, refreshed every minute, and is hosted outside the infrastructure it watches. Worker counts, per-provider breakdowns and latency are not public: that visibility comes with a dedicated endpoint.

## Related

* [Privacy Policy (Inference)](/privacy-policy-inference.md) · [Terms of Service (Inference)](/terms-of-service-inference.md)
* [Models and pricing](/inference-api/models-and-pricing.md): provider pinning in full
* [Inference API quickstart](/inference-api/inference-api.md): base URL, keys, first request
