Choosing and managing models
Install compatible models, select a conversation model and diagnose loading or memory issues.
In this topic
A model supplies the reasoning or generation behind a Mellow conversation. Your agent adds its instructions and enabled capabilities; the selected device determines where the work runs. Keep those choices separate when setting up a new chat.
Use Local Models to manage downloaded language models and Providers to connect model services. Speech and image models have their own settings under Voice and Images.
What you can run
| Route | Setup | Main consideration |
|---|---|---|
| Downloaded local language model | Install a compatible model in Local Models | Memory use, supported architecture and capabilities |
| Apple Foundation | Eligible macOS and Apple Intelligence configuration | System availability and a smaller context budget |
| Configured provider | Endpoint and its required account or key | Network availability, provider capabilities and account limits |
| Paired host | A connected Mac with its own ready model | Work and model execution belong to that host |
Cloud workspace agents are a separate discovery and task surface. Connecting a workspace does not install local models or configure every model provider.
Get started
- Open Local Models and inspect the available models and their size estimates.
- Choose a compatible model that leaves room for macOS and your other apps.
- Start its download and wait until the app reports it installed.
- Open a new conversation and choose the model in the composer.
- Send a short text-only request before testing files, images or tools.
A complete download and a successful first response are different checks. A model can exist on disk but fail to load because of memory pressure, missing files or an unsupported architecture.
Choosing a model in chat
The picker identifies the provider and model. Select the provider first when more than one offers similarly named models. Options such as reasoning effort appear only where the selected route supports them.
Before a longer task, confirm the model shown in the composer and the device running the conversation. An agent's saved preference is a starting point; a conversation may have a different selection.
Model options
| Option | Effect | Useful practice |
|---|---|---|
| Context limit | Bounds the text available to a request | Leave room for instructions, tools and the response |
| Output limit | Bounds generated text | Increase only when the task needs a longer answer |
| Temperature | Adjusts sampling variation | Keep other settings constant when comparing answers |
| Reasoning effort | Selects a provider-supported reasoning mode | Use only values offered for that model |
| Thinking controls | Expose supported reasoning behavior | Do not assume all local models implement them |
Provider controls are not interchangeable. An OpenAI-compatible endpoint may accept basic text requests while rejecting a particular reasoning or multimodal option.
Local models (MLX)
Mellow discovers compatible local model bundles and supports repository-based downloads. Compatibility depends on the model architecture and its packaged files, not just its file extension or a model name containing “MLX”.
Downloading
Use the model's download controls to pause, resume or cancel transfers. Check available disk space for the complete bundle. If repository authentication is required, resolve that access before repeatedly restarting the transfer.
Which model should I pick?
Begin with a smaller supported model and the task you actually need: concise writing, code explanation or tool use. Compare complete results rather than only the first-token speed. A model that handles plain text well may not support images or reliable structured tool calls.
How much memory does a model need?
Download size is not the full runtime budget. Model weights, attention cache, conversation length and simultaneous work all consume memory. Quantization can reduce weight size, but does not eliminate context growth or other app memory usage.
If the Mac becomes unresponsive, stop the run, shorten the context or choose a smaller model. Loading multiple large models is not a substitute for checking available memory.
Models that understand images and audio
Use capabilities reported by the installed model and supported runtime. Attaching an image to a text-only model does not give it visual understanding. Voice transcription can turn speech into text independently of whether the chat model accepts audio directly.
Updates, verification, and repair
Refresh the model's state before diagnosing a download that looks incomplete. Use the available verification or repair controls for that bundle. Keep a working model available while trying an updated one, especially if it is used by a routine.
Deleting a model frees its files but does not select a replacement for every agent or saved conversation. Review defaults that referenced it.
Open a model from Hugging Face
Import a compatible repository using its actual publisher/repository ID. Inspect what the importer recognizes before downloading. Do not substitute a Mellow-branded publisher into a third-party model ID; repository identifiers must match the source of the weights.
Where models live
The app manages a model directory and can discover existing compatible bundles, including supported external paths. Consult the installed configuration before moving large downloads. See Storage for data locations and Configuration for overrides.
Keep external drives mounted while a model from that drive is in use. A catalog entry can remain visible after its files become unavailable.
Apple Foundation Models
Foundation uses Apple's on-device framework when the Mac and operating system report it available. It has separate OS requirements from the Mellow app. See Apple's on-device model for setup, availability states and API discovery.
Cloud providers
Configure your own account under Providers, refresh its models, and test a small request. A model listed in a screenshot is not necessarily available to your account. Provider connections explains endpoint types, credentials and recovery.
Image models
Image generation, editing and upscaling use operation-specific models. Manage them under Images and select a model with the capability required for the job. See Images and video.
Troubleshooting
| Symptom | Check first |
|---|---|
| Model not found | Exact selected ID, installed bundle and connected provider |
| Download stops | Disk space, network and repository access |
| Bundle exists but cannot load | Completeness, supported architecture and memory |
| Response becomes slow in a long chat | Context size and memory pressure |
| Tools do not work | Model tool support and agent capability configuration |
| Images are ignored | Vision support for this model and route |
| Provider model is missing | Authentication, discovery and account entitlement |
Under the hood
The local server's model discovery exposes the IDs clients should request. Use discovery instead of hard-coding a sample model name. A successful list operation establishes catalog access; follow it with a small completion to check inference.
When writing a client, distinguish model-not-found, loading and generation failures. Retry a transient connection error differently from an unsupported model or missing credential. See Inference runtime and HTTP API for the request lifecycle.
Continue exploring · Models, voice and mediaProvider connections →Configure endpoints, protocol types and credentials for your model accounts.