BYOK

Your OpenAI, Anthropic, or Gemini key. Your invoice. Your DPA.

On Enterprise, BeforeQuery calls the model on your provider account with your key. You keep the model choice, the usage bill, and the data-processing agreement you already have with OpenAI / Anthropic / Google. BeforeQuery is the platform layer; the LLM traffic runs on you.

Capabilities

How BYOK works

Per-workspace credentials, per-knowledge-base overrides — every customer-facing answer runs on the account you configured

01
OpenAI, Anthropic, Gemini

Bring Your Provider Credential

  • Configure a provider credential per workspace
  • Optionally override per knowledge base — different KBs on different providers
  • Credentials encrypted at rest with per-workspace envelope keys
  • Rotate or revoke in the dashboard; no re-deploy required
  • Fallback to the deployment provider when a credential is missing
02
Everywhere the model is called

Coverage Across the Answer Path

  • Streaming widget and portal answers — on your credential
  • Multi-knowledge-base answers — on your credential
  • Slack / Discord / Teams bots (AnswerForBot) — on your credential
  • Form deflector, including grounding verification — on your credential
  • Reranking on the LLM grader fallback — on your credential
03
Honest about the gaps

What Still Runs on Deployment Provider

  • Internal diagnostic and eval tooling (deployment models, so scores are comparable across customers)
  • Conversation title generation (small fixed-cost call)
  • Image-to-text (DescribeImage) for chat attachments
  • Query planning (SearchService's own provider today)
  • Hosted answer-quality re-scoringers (Cohere/Voyage/Jina — not LLM providers)
Why BYOK exists

The two questions Enterprise buyers actually ask

'Whose account is charged for the tokens?' and 'Whose DPA covers the data?' — the two questions every enterprise procurement conversation lands on. BYOK answers both with 'yours'. Traffic runs on your OpenAI / Anthropic / Gemini account under the terms you already negotiated. BeforeQuery is the platform above the model, not the LLM vendor.

  • Volume LLM contracts you already negotiated apply directly
  • Model choice is yours — no vendor lock-in via the platform
  • Data-processing terms are yours — no additional third-party review
  • Rotate the credential without re-deploying anything
  • Fallback to deployment provider keeps the workspace up if the credential fails
What teams say

Deployed in production.

“Deployed on our docs site in an afternoon. Every answer shows the source, and the abstention gate means we've never had a customer complain about a made-up answer.”
Priya Nair
Head of Customer Support · Supabase
“The knowledge base connected to our Slack, Confluence, and helpdesk in one setup. On-call teams get the same grounded answer whether they ask in chat, in the widget, or from Cursor.”
Tom Richter
IT Operations Manager · Grafana Labs
“The gap analytics turned into a real docs backlog. Deflection went up because we finally knew which pages were missing — the AI told us.”
Ana Castillo
VP of Customer Experience · Clerk

Frequently Asked Questions

Common questions about Bring Your Own Model (BYOK)

OpenAI, Anthropic, and Google Gemini for the chat model. Cross-encoder answer-quality scorers (Cohere, Voyage, Jina) are supported separately as hosted services — they're not LLMs, so BYOK for them is 'use your hosted-answer-quality scorer key' rather than a full model swap.
Yes. Per-workspace credential is the default; per-knowledge-base override is available for the case where you want, for example, Claude Sonnet for a customer-facing KB and GPT-4o for internal ops. The router picks the closest override at call time.
The call fails and the user sees a friendly 'try again' message. There's a configurable fallback to the deployment provider so the workspace stays up during transient credential issues — you can turn it off if strict single-provider is a requirement.
The key is encrypted at rest with a per-workspace envelope key. Only the request path uses it, and only for the calls that route through the workspace's credential. It's not visible to other workspaces and not used for training or logging beyond the invocation record.
BeforeQuery is priced for the platform (retrieval, connectors, evals, analytics, surfaces). BYOK removes the LLM inference from our cost basis, so Enterprise deals with BYOK typically settle on a lower platform fee. Non-BYOK plans (Free, Pro) run on the deployment provider and include LLM inference in the plan price.
No — BYOK is an Enterprise feature today. Free and Pro run on the deployment provider so the pricing stays flat and self-serve. Enterprise adds BYOK, SSO, DPA, regional residency, and volume terms.

Run the platform, keep the LLM contract

Talk to us about Enterprise — BYOK, DPA, regional residency, and SSO all move together on Enterprise plans.

Get Started Free