Raw server access
Ta treść nie jest jeszcze dostępna w Twoim języku.
Raw server access
Section titled “Raw server access”LM Nexus can start and manage llama-server without forcing every request through an extra Nexus proxy.
You have two paths.
Integrated LM Nexus /v1
Section titled “Integrated LM Nexus /v1”Model Manager exposes an OpenAI-compatible /v1 API.
When you use it, LM Nexus can resolve the requested managed profile, start it when appropriate, forward the request to the managed server, and keep the request tied into Model Manager’s runtime behavior.
Direct llama.cpp endpoint
Section titled “Direct llama.cpp endpoint”LM Nexus also shows the port of the underlying llama-server.
If a client connects directly to that port, the request goes straight to llama.cpp:
Integrated:client -> LM Nexus /v1 -> managed llama-server
Direct:client ----------------> managed llama-serverFor example:
curl http://127.0.0.1:PORT/v1/chat/completionsReplace PORT with the actual runtime port shown by LM Nexus.
When direct access is useful
Section titled “When direct access is useful”Use the direct endpoint when you want to:
- connect another OpenAI-compatible client straight to llama.cpp;
- test the runtime without the LM Nexus API in the middle;
- check compatibility at the llama.cpp boundary;
- let another tool own the inference client while LM Nexus only manages the server.
Use the integrated /v1 path when you want the request to go through LM Nexus Model Manager behavior.
LM Nexus still owns the process
Section titled “LM Nexus still owns the process”Using the direct endpoint does not make the server external.
If LM Nexus started that llama-server, LM Nexus should still be the thing that stops it.