Skip to main content
Enclave
Model integrations

Use Groq in Enclave (ultra-low-latency inference)

Enclave supports Groq via BYOK: enter your Groq API key to use its hosted open-source models (such as the Llama family) as the base for Enclave characters. Groq is known for extremely low inference latency, ideal for latency-sensitive real-time chat. Enclave doesn't resell quota and bills at Groq's official rates.

Last reviewed:

The short answer

Enter your Groq API key and Enclave's characters reply with Groq's ultra-low-latency inference—closer to an instant reply.

Key points at a glance

  • BYOK: use your own Groq account and quota.
  • Ultra-low-latency inference, ideal for real-time, high-frequency conversational characters.
  • Mix with other providers: route speed-critical characters to Groq and others elsewhere.

How to connect

In model settings choose Groq (an OpenAI-compatible provider), enter your API key, point the base URL at Groq's endpoint, then set it as default or assign per character. Enclave lets multiple providers coexist, so you can move just the latency-sensitive characters to Groq.

Who it's for

Users who value response speed and want chat to feel smooth—especially latency-sensitive scenarios like voice companionship and real-time group chat. With Enclave's memory and proactivity, you get both speed and being remembered.

Key facts

Integration
BYOK (OpenAI-compatible endpoint)
Billing
Billed at Groq's official rates; no Enclave markup
Highlight
Ultra-low inference latency
Data flow
Direct from your instance when self-hosted

Ready to try it?

Open it in your browser — no credit card, no install.

Related

Connect other models

Use Groq in Enclave (ultra-low-latency inference) · Enclave