Hosted inference
Distinguish model-service readiness and billing from Cloud workspace sign-in.
In this topic
Mellow can use a hosted inference service when that service is configured and available. This is separate from signing in to a Cloud workspace, which supplies workspace-scoped agents and tasks. A successful workspace sign-in does not establish a hosted model catalog or a credit balance.
Use this distinction when a Cloud workspace is connected but the model browser is empty. The workspace may be ready for agent discovery while inference needs a different configured service.
Choose the connection you need
| Your goal | Configuration to inspect |
|---|---|
| Run a downloaded model on this Mac | Local Models |
| Use your own hosted model account | Providers |
| Discover published agents in an organization | Mellow Cloud, then the authenticated Cloud workspace |
| Use managed hosted inference | Its model service, identity and account availability |
| Run work on another paired Mac | Devices, then that host's agent and provider |
For direct provider accounts, follow Provider connections. For workspace tasks, use Cloud workspaces.
Get started
First inspect the service status in the installed app. The current Router configuration is disabled unless explicitly enabled. Its presence in source code or a settings route is not evidence that a public service has been provisioned.
When a deployment supplies hosted inference, enable it through its account controls, complete any required identity setup, and refresh the model catalog. Select an advertised model and send a small request. Confirm that it produces an answer before relying on it for a longer task.
Do not purchase credits as a remedy for a missing service endpoint or failed model discovery. Resolve availability and account errors first.
Finding models
A usable catalog should identify the model and expose the controls it supports. Choose a model from that catalog rather than typing an example name from a screenshot. Provider model names and availability can change independently of the Mellow app version.
An empty catalog can mean discovery failed, the service is disabled, or your account has no eligible models. It is not sufficient evidence that you have exhausted a balance. Keep the error visible and distinguish retrying discovery from signing in again.
Credits
Where a configured service supports metered inference, review its current balance, prices and billing scope in that service's account controls. This reference does not promise welcome credits, a conversion rate or an included subscription.
A workspace role, provider subscription and hosted-inference balance are different things. Check which account would be charged before starting a large request. Hosted image and video operations can have separate quote and approval steps; see Images and video.
Checking your balance in chat
If the installed build displays a balance or usage control, treat it as information for the service named there. A cached balance may lag a recent operation. Refresh through the service's controls before drawing conclusions about a failed request.
Workspaces and shared credits
Use the account and billing information returned by the configured workspace. Do not assume that a published workspace agent shares a personal inference balance. Workspace permissions govern which actions and agents are available; billing depends on that deployment's contract.
Turning it off
Disabling managed inference removes that route from use; local models and independently configured providers remain separate choices. If a saved conversation selected an unavailable hosted model, select another ready model before retrying.
Your privacy
Hosted inference sends the request content required by the selected model to its service. Review both the service policy and any tools enabled for the conversation. Local storage of chat history does not mean remote inference received no content.
Billing diagnostics and conversation content are different records. Review an export before sharing it, and avoid treating the existence of a local billing ledger as proof of a remote service's retention behavior.
Troubleshooting
| State | Meaning to investigate | Next action |
|---|---|---|
| Router is not configured | Build/deployment does not supply usable managed inference | Use a ready local model or your configured provider |
| Workspace connected, no models | Workspace connection and inference are separate | Check the provider/model route you selected |
| Model catalog fails | Discovery, service or account issue | Read the actual error and retry discovery after correcting it |
| Request reports insufficient balance | The selected service rejected metered work | Review that account's balance and billing scope |
| Model disappeared | Catalog or permissions changed | Refresh and select an advertised model |
| Request finished without an answer | Generation or rendering failed | Inspect the run error before retrying a billable task |
Under the hood
The Router integration has its own enabled state and identity-signed requests. Keyless loopback spending is a separate opt-in and defaults off. A local process should not be given billing authority merely because it can reach the local server.
Release configuration fixes the production service address; debug configuration supports controlled test overrides. These implementation choices do not certify deployment availability. Validate the actual account, catalog and a small request in the environment being used.
Continue exploring · Models, voice and mediaSpeaking and dictation →Configure recognition, chat input, text insertion, wake phrases and stop behavior.