Case Study
SupportAI — a drop-in AI support widget that knows when to hand off
SupportAI answers product questions from a real FAQ knowledge base over pgvector, streams the answer into a chat widget, and escalates to a human after two consecutive misses. The non-trivial part is not any one piece. It is retrieval, confidence, handoff, and admin tooling all working together.
- Role
- Solo — design, engineering, documentation.
- Stack
- Next.js 15 · Supabase pgvector · Groq / Gemini · Vercel.
- Timeline
- Built over 2 days.

The problem
A small SaaS does not need a chatbot that guesses. It needs something narrower: answer the questions the docs already cover, word for word, and the moment a question falls outside the docs, get a human involved with context attached.
That sounds simple until the pieces have to agree. Retrieval has to find the right rows. The model has to trust the retrieved context and nothing else. Something has to decide, per turn, whether the answer counts — and count consecutive failures without losing state across reloads. Then an admin needs to see what happened.
SupportAI is that whole loop: a floating widget, a streaming RAG API, a two-strike handoff counter, and a small admin panel for FAQs and conversation logs.
The architecture
One request flows through eleven steps. Embeddings come from Gemini, generation defaults to Groq with a Gemini fallback, and every turn is persisted to Postgres before the stream closes.


Technical decisions
Tool-call handoff, not similarity threshold
I started with a threshold: zero matches below 0.5 means no answer. Then I measured. Junk queries scored 0.57 to 0.59, real matches 0.68 to 0.77. A 0.04 to 0.09 gap is not a threshold, it is a coin flip. So I stopped thresholding and gave the model a reportNoAnswer tool to call instead of guessing. The handoff signal is now deterministic: the model either called the tool or it did not.
junk 0.57-0.59 vs real 0.68-0.77 // too thin to threshold
Server-authoritative fallback copy
When the model has nothing to say, I do not let it write the apology. Model-generated fallbacks drift between turns and occasionally promise things. One fixed string is persisted and streamed identically on every no-answer turn. It costs nothing and reads the same every time.
FALLBACK_TEXT = "I don't have that in my FAQs..."
PostgREST silently nulls vector RPC args
The first version of the similarity search returned zero rows with no error. The cause: passing a JS number[] to a vector-typed RPC argument, which PostgREST silently drops instead of coercing. The fix is a text argument with an internal cast, and the client sends the pgvector text format.
query_embedding: JSON.stringify(embedding) // "[0.02,-0.01,...]"
RLS default-deny reads as empty, not forbidden
Supabase enables row-level security by default, and an unauthorized read returns an empty table with no error. The FAQ table looked seeded and empty at the same time. One public read policy on faqs fixed it. The other three tables stay service_role-only by design.
faqs_public_read // SELECT for anon; nothing else
Switched providers mid-build without touching the route
Gemini 2.0-flash was retired mid-project. Its replacement hit a 20-request-per-day free-tier ceiling I measured, not one I read about. Because the route talks to the AI SDK and not to Google, the fix was changing one model ID, then flipping primary to Groq gpt-oss-120b at 664ms per turn with no daily cap.
groq("openai/gpt-oss-120b") // 664ms, was 54s+ on GeminiDisabled Gemini thinking mode
Before the provider switch, single turns took up to 45 seconds. The embedding plus FAQ context already contains everything needed for these answers, so token-by-token reasoning added nothing. Setting the thinking budget to zero brought turns under 2 seconds.
thinkingConfig: { thinkingBudget: 0 } // 45s -> 1.5sWhat I would do differently
- Mid-stream provider fallback. Today only a synchronous failure fails over; a 404 halfway through a stream loses the turn. The stream needs a catch point that can restart on the fallback model.
- Real rate limiting on /api/handoff. It currently relies on the platform edge. Upstash Redis with a sliding window is the intended fix.
- Admin auth is a shared password, not a user system. Fine for a demo, wrong for multi-tenant.
Under the hood


- Codebase
- ~3,400 lines across ~70 files, 8 commits
- Verify
- npm run seed · tsc --noEmit · npm run build, all green
- Fallback order
- Groq gpt-oss-120b, then Gemini 3.1-flash-lite
- Cost
- Zero paid tools: Gemini, Groq, Supabase, Vercel, Resend, GitHub free tiers