← Selected work
Case Study · AI Engineering · Solo Build

EasyTab AI

A 24/7 AI operations manager for small civil-construction contractors.

Every business gets a private “Brain” — a retrieval-augmented AI that learns their projects, clients, invoices, and documents, then does their office work: answering questions, chasing overdue payments, drafting bid proposals, replying to email, and finding new government tenders. Built for owner-operators who run their company from a phone and a Gmail inbox.

Status
Pre-launch
Role
Founder, solo build
Built with
Multi-tenant RAG · Next.js 16 + FastAPI · Claude API · pgvector · 6 external APIs · 9 E2E test suites

Built solo and pre-launch, tested on seeded demo firms. Figures come from code and test runs, not projections.

01 — The Problem

Small contractors lose money to office work, not construction work

The owner of a 2–15 person contracting business is also the estimator, the collections department, and the receptionist: chasing overdue invoices by hand, writing bid proposals from scratch against dense tender documents, answering “what did we quote that client last year?” by digging through folders, and missing calls — and therefore jobs — while on site.

Generic CRMs don’t help, because the knowledge lives in unstructured documents: contracts, tenders, old proposals that no form-based tool can read. EasyTab’s bet is that a per-company RAG system over exactly those documents can automate the drafting work — while keeping a human on every send button.

02 — What It Does

One Brain, six jobs

● live — verified end-to-end in a real browser against real APIs  ·  ● pending — code-complete and tested, awaiting external account credentials

Live-verified

Ask-anything chat

Answers business questions from the company’s own data, with sources. Every dashboard write — new invoice, client, project — is re-ingested so answers stay current.

Live-verified

Document onboarding

Upload contracts and PDFs at signup; async ingestion with live per-file progress: queued → processing → learned.

Live-verified

Payment reminder drafting

Unpaid invoices auto-tiered by days overdue (upcoming → polite → firm → urgent). AI drafts the chase message in the company’s tone, citing the real invoice number — human review before every send.

Live-verified

Bid / proposal generator

Upload a tender PDF; the AI writes a structured proposal citing only real past projects from the Brain — and leaves pricing explicitly blank rather than inventing a number.

Live-verified

Email responder

Gmail OAuth integration verified against the live Gmail API: connect, read, classify (RFQ / invoice / complaint…), draft a grounded reply, edit, send.

Live-verified

Multi-tenant auth & dashboard

Email confirmation, password reset, rate limiting, per-owner data isolation. KPI tiles, audit log, global search, detail pages, dark mode.

Pending credentials

Missed-call AI voice bot

Telnyx Call Control: signature-verified webhooks, grounded conversation loop capped at 8 exchanges, transcript + AI summary to the owner. Blocked on carrier provisioning, not code.

Pending credentials

AI lead finder

Searches SAM.gov federal Contract Opportunities across civil-construction NAICS codes; the LLM scores each result’s fit.

Pending credentials

WhatsApp send & Stripe billing

Meta Cloud API reminders; Stripe Checkout, Customer Portal, and signature-verified webhooks syncing entitlements with a 14-day in-app trial. Each degrades gracefully to a clear message — never a crash.

03 — Architecture

Two services, one Brain

The web app is the only user-facing surface; all AI lives behind it. No cron jobs — everything is event- or request-driven.

System flow — drawn at the boxes-and-arrows level

  1. 1

    Next.js 16 web app

    App Router + Server Actions. Owns auth, multi-tenant authorization, the dashboard, and all third-party webhooks (Stripe, telephony). Per-user authorization happens entirely in this layer.

  2. 2

    Python FastAPI service — “the Brain”

    Called server-to-server with a shared-secret header. Has no user concept: it only ever sees company-scoped requests the web layer has already authorized.

  3. 3

    RAG pipeline → Postgres + pgvector

    Parse → chunk (~500 words) → embed locally (BGE-small, 384-dim) → store in Supabase Postgres. Ingestion is async via background tasks; a job-status table drives the live progress UI.

  4. 4

    Claude generates everything

    Chat answers, reminder drafts, proposals, email replies, voice-call turns, lead scoring — all Anthropic API calls, made only from the Brain. Every prompt carries an explicit anti-invention instruction.

  5. 5

    Event-driven integrations

    Stripe and Telnyx push signature-verified webhooks; Gmail (raw REST + OAuth refresh), WhatsApp (Meta Cloud API), and SAM.gov are called on demand. Missed call: webhook → verify → grounded reply → speak & listen → summarize → notify owner.

  6. 6

    Deployment-ready

    Dockerized Brain for a PaaS, web app for Vercel, managed Postgres, step-by-step runbook. Webhook development runs through a Cloudflare tunnel.

04 — Technical Decisions

Choices, alternatives, and why

Local embeddings instead of an embedding API

fastembed BGE-small runs in-process: free, no vendor dependency, no per-token cost on every document. The tradeoff — CPU-bound ingestion on huge files — was solved with async background jobs, not bigger hardware.

Why — LLM spend is reserved for generation, where quality actually shows.

Exact vector scan, no ANN index — on purpose

An ivfflat index silently returned zero rows on small per-company datasets. At hundreds-to-thousands of chunks per tenant, exact nearest-neighbor is both correct and fast; the index is deferred until a tenant nears ~50k chunks.

Why — correctness over premature optimization.

Guardrails against invented output, enforced in tests

Every call is grounded in retrieved company data and instructed never to invent figures. The proposal generator leaves [Pricing to be inserted] rather than guessing — and an E2E test asserts the blank stays blank. After the AI once fabricated an invoice number, the real reference is now passed into the prompt.

Why — a “no fictional data” rule covers code, docs, and demos alike.

Human-in-the-loop on everything outbound

Reminders, email replies, and WhatsApp sends are always draft → review → send.

Why — a wrong amount in a money-related message to a real client is worse than the friction of one click.

Hand-rolled REST clients by default, SDKs only where security demands

Gmail, WhatsApp, and SAM.gov are plain fetch calls. Stripe and Telnyx got SDKs specifically for webhook signature verification — HMAC schemes with timestamp tolerance are easy to get subtly wrong, and getting them wrong means anyone can forge a “subscription active” event or a fake call.

Why — minimal dependencies, except where a mistake becomes a security hole.

App-layer tenant isolation

Every query is scoped by company, and companies by authenticated owner, via two required helpers — simpler and more auditable at this scale than per-request RLS, with RLS kept on as a deny-all backstop.

Why — auditability first, defense in depth second.

05 — Hardest Problem

The retrieval bug that returned 200 OK

The RAG chat started confidently answering “I don’t have that information” — for a company whose data was verifiably in the database. No errors anywhere: ingestion succeeded, embeddings existed, the API returned 200s. The system was silently wrong, which for a trust-the-answer product is the worst failure mode.

Diagnosis meant walking the pipeline backwards with the same question at each layer. The LLM was fine — given chunks, it answered correctly. The query embedding was fine. But the vector similarity search returned zero rows for questions that should have matched. The culprit: the ivfflat ANN index partitions vectors into clusters and probes only a few — which works at scale, but on a tenant with a handful of chunks it can probe only empty clusters and return nothing, with no error. The index was tuned for a scale the data didn’t have.

The fix: drop the index, run exact scans, encode the decision in the schema with a comment explaining when to reinstate it (~50k chunks per tenant), and re-verify in a real browser that every previously-failing question now retrieved and answered.

An optimization that degrades silently is worse than no optimization — and retrieval quality must be tested end-to-end with real data. A 200 response proves nothing.

06 — Engineering Metrics

Measured, not marketed

45s → 1.5s
onboarding upload response, after moving ingestion to async background jobs
$0
embedding cost — local model; LLM spend is generation only
10+
real bugs caught by browser-level E2E tests before any user saw them
9
Playwright E2E suites, full run green — incl. cross-tenant isolation
~14s
chat round-trip; proposals in 30–60s (measured, unoptimized)
6
external APIs integrated — two of them signature-verified webhook flows

Pre-launch product — these are engineering measurements, not customer claims. Planned pricing: $497–$997/mo subscription tiers, already built into the Stripe integration.

07 — Tech Stack

What it’s built with

Frontend

  • Next.js 16 (App Router, Server Actions)
  • React 19 · TypeScript
  • Tailwind CSS v4 · dark mode

Backend

  • Python FastAPI (RAG service)
  • Next.js server actions / API routes
  • Supabase — Postgres + pgvector + Auth

AI / LLM

  • Claude (Anthropic API) — all generation
  • fastembed BGE-small local embeddings
  • Custom chunking & retrieval — no framework

Integrations

  • Gmail REST (OAuth) — live
  • Stripe · WhatsApp Cloud · Telnyx · SAM.gov
  • Resend for auth email

Infra & Testing

  • Docker + Vercel/Render, deploy runbook
  • Cloudflare tunnel for webhook dev
  • 9 Playwright E2E suites · strict TS · ESLint
08 — My Role

Solo, end-to-end

I designed the product and architecture, built both services — the Next.js frontend, auth, and billing, and the Python RAG service — wrote the retrieval pipeline and every prompt, integrated six external APIs including two webhook-driven ones, set up multi-tenant auth, and wrote the Playwright E2E suite that verifies every feature in a real browser.

I used AI coding agents as a force multiplier throughout, with a working method built around them: a living handoff document updated after every task, browser-verified acceptance for every feature, and hard guardrails — no invented data, human review on all outbound messages — enforced in code and tests.

09 — What’s Next

Roadmap

Activate pending integrations and deploy

The remaining work is account setup, not code: WhatsApp, Stripe, SAM.gov, and the voice bot each need a credential only the business owner can obtain — a deliberate pattern where every integration degrades gracefully without its key, so code could ship ahead of credentials.

Proactive signals over reactive answers

A five-stage validation study on public NYC construction permit data for a lead-signal engine — including rebuilding a scoring rubric that initially produced only 2 distinct scores across an entire feed (420 after the rebuild). Lesson: validate data before designing against it.

Harden for scale

Distributed rate limiting, WhatsApp template messaging for first-contact reminders, and the ANN index once tenant data justifies it.

Next case study
FieldLog AI →

Want the full walkthrough? I’ll demo it live.

Get in touch