EasyTab AI
A 24/7 AI operations manager for small civil-construction contractors.
Every business gets a private “Brain” — a retrieval-augmented AI that learns their projects, clients, invoices, and documents, then does their office work: answering questions, chasing overdue payments, drafting bid proposals, replying to email, and finding new government tenders. Built for owner-operators who run their company from a phone and a Gmail inbox.
Built solo and pre-launch, tested on seeded demo firms. Figures come from code and test runs, not projections.
Small contractors lose money to office work, not construction work
The owner of a 2–15 person contracting business is also the estimator, the collections department, and the receptionist: chasing overdue invoices by hand, writing bid proposals from scratch against dense tender documents, answering “what did we quote that client last year?” by digging through folders, and missing calls — and therefore jobs — while on site.
Generic CRMs don’t help, because the knowledge lives in unstructured documents: contracts, tenders, old proposals that no form-based tool can read. EasyTab’s bet is that a per-company RAG system over exactly those documents can automate the drafting work — while keeping a human on every send button.
One Brain, six jobs
● live — verified end-to-end in a real browser against real APIs · ● pending — code-complete and tested, awaiting external account credentials
Ask-anything chat
Answers business questions from the company’s own data, with sources. Every dashboard write — new invoice, client, project — is re-ingested so answers stay current.
Document onboarding
Upload contracts and PDFs at signup; async ingestion with live per-file progress: queued → processing → learned.
Payment reminder drafting
Unpaid invoices auto-tiered by days overdue (upcoming → polite → firm → urgent). AI drafts the chase message in the company’s tone, citing the real invoice number — human review before every send.
Bid / proposal generator
Upload a tender PDF; the AI writes a structured proposal citing only real past projects from the Brain — and leaves pricing explicitly blank rather than inventing a number.
Email responder
Gmail OAuth integration verified against the live Gmail API: connect, read, classify (RFQ / invoice / complaint…), draft a grounded reply, edit, send.
Multi-tenant auth & dashboard
Email confirmation, password reset, rate limiting, per-owner data isolation. KPI tiles, audit log, global search, detail pages, dark mode.
Missed-call AI voice bot
Telnyx Call Control: signature-verified webhooks, grounded conversation loop capped at 8 exchanges, transcript + AI summary to the owner. Blocked on carrier provisioning, not code.
AI lead finder
Searches SAM.gov federal Contract Opportunities across civil-construction NAICS codes; the LLM scores each result’s fit.
WhatsApp send & Stripe billing
Meta Cloud API reminders; Stripe Checkout, Customer Portal, and signature-verified webhooks syncing entitlements with a 14-day in-app trial. Each degrades gracefully to a clear message — never a crash.
Two services, one Brain
The web app is the only user-facing surface; all AI lives behind it. No cron jobs — everything is event- or request-driven.
System flow — drawn at the boxes-and-arrows level
-
1
Next.js 16 web app
App Router + Server Actions. Owns auth, multi-tenant authorization, the dashboard, and all third-party webhooks (Stripe, telephony). Per-user authorization happens entirely in this layer.
-
2
Python FastAPI service — “the Brain”
Called server-to-server with a shared-secret header. Has no user concept: it only ever sees company-scoped requests the web layer has already authorized.
-
3
RAG pipeline → Postgres + pgvector
Parse → chunk (~500 words) → embed locally (BGE-small, 384-dim) → store in Supabase Postgres. Ingestion is async via background tasks; a job-status table drives the live progress UI.
-
4
Claude generates everything
Chat answers, reminder drafts, proposals, email replies, voice-call turns, lead scoring — all Anthropic API calls, made only from the Brain. Every prompt carries an explicit anti-invention instruction.
-
5
Event-driven integrations
Stripe and Telnyx push
signature-verifiedwebhooks; Gmail (raw REST + OAuth refresh), WhatsApp (Meta Cloud API), and SAM.gov are called on demand. Missed call: webhook → verify → grounded reply → speak & listen → summarize → notify owner. -
6
Deployment-ready
Dockerized Brain for a PaaS, web app for Vercel, managed Postgres, step-by-step runbook. Webhook development runs through a Cloudflare tunnel.
Choices, alternatives, and why
Local embeddings instead of an embedding API
fastembed BGE-small runs in-process: free, no vendor dependency, no per-token cost on every document. The tradeoff — CPU-bound ingestion on huge files — was solved with async background jobs, not bigger hardware.
Why — LLM spend is reserved for generation, where quality actually shows.
Exact vector scan, no ANN index — on purpose
An ivfflat index silently returned zero rows on small per-company datasets. At hundreds-to-thousands of chunks per tenant, exact nearest-neighbor is both correct and fast; the index is deferred until a tenant nears ~50k chunks.
Why — correctness over premature optimization.
Guardrails against invented output, enforced in tests
Every call is grounded in retrieved company data and instructed never to invent figures. The proposal generator leaves [Pricing to be inserted] rather than guessing — and an E2E test asserts the blank stays blank. After the AI once fabricated an invoice number, the real reference is now passed into the prompt.
Why — a “no fictional data” rule covers code, docs, and demos alike.
Human-in-the-loop on everything outbound
Reminders, email replies, and WhatsApp sends are always draft → review → send.
Why — a wrong amount in a money-related message to a real client is worse than the friction of one click.
Hand-rolled REST clients by default, SDKs only where security demands
Gmail, WhatsApp, and SAM.gov are plain fetch calls. Stripe and Telnyx got SDKs specifically for webhook signature verification — HMAC schemes with timestamp tolerance are easy to get subtly wrong, and getting them wrong means anyone can forge a “subscription active” event or a fake call.
Why — minimal dependencies, except where a mistake becomes a security hole.
App-layer tenant isolation
Every query is scoped by company, and companies by authenticated owner, via two required helpers — simpler and more auditable at this scale than per-request RLS, with RLS kept on as a deny-all backstop.
Why — auditability first, defense in depth second.
The retrieval bug that returned 200 OK
The RAG chat started confidently answering “I don’t have that information” — for a company whose data was verifiably in the database. No errors anywhere: ingestion succeeded, embeddings existed, the API returned 200s. The system was silently wrong, which for a trust-the-answer product is the worst failure mode.
Diagnosis meant walking the pipeline backwards with the same question at each layer. The LLM was fine — given chunks, it answered correctly. The query embedding was fine. But the vector similarity search returned zero rows for questions that should have matched. The culprit: the ivfflat ANN index partitions vectors into clusters and probes only a few — which works at scale, but on a tenant with a handful of chunks it can probe only empty clusters and return nothing, with no error. The index was tuned for a scale the data didn’t have.
The fix: drop the index, run exact scans, encode the decision in the schema with a comment explaining when to reinstate it (~50k chunks per tenant), and re-verify in a real browser that every previously-failing question now retrieved and answered.
An optimization that degrades silently is worse than no optimization — and retrieval quality must be tested end-to-end with real data. A 200 response proves nothing.
Measured, not marketed
Pre-launch product — these are engineering measurements, not customer claims. Planned pricing: $497–$997/mo subscription tiers, already built into the Stripe integration.
What it’s built with
Frontend
- Next.js 16 (App Router, Server Actions)
- React 19 · TypeScript
- Tailwind CSS v4 · dark mode
Backend
- Python FastAPI (RAG service)
- Next.js server actions / API routes
- Supabase — Postgres + pgvector + Auth
AI / LLM
- Claude (Anthropic API) — all generation
- fastembed BGE-small local embeddings
- Custom chunking & retrieval — no framework
Integrations
- Gmail REST (OAuth) — live
- Stripe · WhatsApp Cloud · Telnyx · SAM.gov
- Resend for auth email
Infra & Testing
- Docker + Vercel/Render, deploy runbook
- Cloudflare tunnel for webhook dev
- 9 Playwright E2E suites · strict TS · ESLint
Solo, end-to-end
I designed the product and architecture, built both services — the Next.js frontend, auth, and billing, and the Python RAG service — wrote the retrieval pipeline and every prompt, integrated six external APIs including two webhook-driven ones, set up multi-tenant auth, and wrote the Playwright E2E suite that verifies every feature in a real browser.
I used AI coding agents as a force multiplier throughout, with a working method built around them: a living handoff document updated after every task, browser-verified acceptance for every feature, and hard guardrails — no invented data, human review on all outbound messages — enforced in code and tests.