Use the local API
Use the local API
Section titled “Use the local API”LM Nexus offers two different local OpenAI-compatible paths:
- the integrated Nexus
/v1API; - the direct endpoint of a running managed
llama-server.
They are useful for different reasons.
Use the integrated Nexus /v1 API
Section titled “Use the integrated Nexus /v1 API”1. Open API Serving
Section titled “1. Open API Serving”In Model Manager, select API Serving.
2. Enable API serving
Section titled “2. Enable API serving”Turn on:
Enable OpenAI-compatible /v1 API serving
API Serving is off by default.
After enabling it, the page checks /v1/health and shows the current API base
URL, running-model count, endpoint list, and copyable examples.
3. Check the endpoint
Section titled “3. Check the endpoint”With the normal local backend address, the simplest smoke request is:
curl http://127.0.0.1:8000/v1/modelsThe API Serving page also provides copyable examples for:
- curl;
- Python OpenAI client;
- Linux/WSL CLI;
- PowerShell CLI;
NEXUS_API_URL.
4. Use a running profile id
Section titled “4. Use a running profile id”For /v1/chat/completions, use a running Model Manager profile’s model_id.
The Servable models section on the API Serving page shows the configured profiles and links back to Profiles.
What the Nexus /v1 path does
Section titled “What the Nexus /v1 path does”The integrated path lets LM Nexus resolve Model Manager profiles and proxy OpenAI-compatible requests to the managed runtime.
Conceptually:
your client | vLM Nexus /v1 | vmanaged llama-serverUse this path when you want the request to participate in Nexus Model Manager serving behavior.
Connect directly to the managed runtime
Section titled “Connect directly to the managed runtime”When a profile is running, Model Manager exposes the assigned port of its
underlying llama-server.
The direct path is:
your client | +-----------------> managed llama-serverA direct request therefore looks like:
curl http://127.0.0.1:PORT/v1/modelsReplace PORT with the assigned port shown by the running profile.
Use direct access when you specifically want to test or use the runtime without
the Nexus /v1 serving layer.
Security note
Section titled “Security note”The Enable API Serving switch controls whether LM Nexus accepts new /v1 requests.
It is not authentication, authorization, bind control, or a firewall.
Directly addressed managed-runtime ports are a separate path and are not protected by that switch.
Treat these endpoints as trusted-local development surfaces unless you have explicitly added an appropriate network/security boundary outside LM Nexus.
If the direct endpoint works but Nexus /v1 does not
Section titled “If the direct endpoint works but Nexus /v1 does not”That strongly suggests the managed runtime itself is healthy and the problem is somewhere in the LM Nexus serving path.
Check:
- API Serving is enabled;
/v1 health statusis Ready;- the requested
model_idmatches a servable profile; - the backend base URL is correct;
- Nexus logs show how the request was resolved.
See Raw server access for the architecture distinction.