Key points at a glance
- Private inference: models run on your own GPU/server.
- OpenAI-compatible API — configured much like connecting a cloud API.
- Suits advanced deployments with high concurrency and latency/cost requirements.
How to connect
First spin up an OpenAI-compatible inference endpoint with vLLM, then in Enclave's model settings choose the OpenAI-compatible provider, point the base URL at your vLLM address, and enter the corresponding key (if any). Then set the model as the default or assign it per character.
Who it's for
Teams and advanced users with GPU resources who want to fully privatize inference, or need stable throughput for multi-user/high-concurrency scenarios. With Enclave self-hosting, models and data both live in your own infrastructure.
Key facts
- Integration
- OpenAI-compatible endpoint (base URL)
- Billing
- Self-hosted, no third-party API fees
- Best for
- High concurrency / private inference
- Data flow
- Stays within your own infrastructure
Ready to try it?
Open it in your browser — no credit card, no install.
Related
Connect other models
- Connect OpenAI (GPT-series) models in Enclave
- Connect Anthropic Claude models in Enclave
- Connect Google Gemini models in Enclave
- Connect DeepSeek models in Enclave
- Connect local Ollama models in Enclave (fully offline capable)
- Use Mistral AI models in Enclave
- Use Groq in Enclave (ultra-low-latency inference)
- Use Together AI in Enclave (cloud-hosted open models)
- Use OpenRouter in Enclave (one key, many models)
- Use self-hosted Hugging Face TGI in Enclave