← Selected work
Case Study · AI Engineering · Solo Build

FieldLog AI

Speak for 90 seconds at the end of the day — the paperwork is done.

An AI-powered field documentation platform for construction field engineers and QA inspectors. Speech becomes structured test data, a day’s activity becomes a professional report, and a failed test auto-drafts its own non-conformance report. Incumbents like Procore ($375/user/month) are digital clipboards; FieldLog uses AI as the backbone.

Status
Pre-launch
Role
Founder, solo build
Built with
Voice → structured data · Next.js 14 + Supabase · Claude + Whisper · Offline-first PWA · Two-way Docs sync · 60-check security suite

Built solo and pre-launch, tested on seeded demo firms. Figures come from code and test runs, not projections.

01 — The Problem

1–2 hours of paperwork after a 10-hour shift

A field engineer ends every day with hours of manual documentation: transcribing density gauge readings from handwritten notes, writing a narrative daily report, filling the client’s Word template by hand, drafting a non-conformance report for every failed test, and re-coloring yesterday’s markups on the site plan PDF in Adobe. This happens after the shift, often on a phone, sometimes with no signal.

At typical billing rates that’s $50–100/day of unbillable time per engineer (est.) — and the error mode is worse than the cost: a reading mis-transcribed into a stamped engineering document. Existing tools digitize the clipboard but still make the human do the writing.

02 — What It Does

AI where it helps, deterministic code where it must be right

● AI — model in the loop, always human-reviewed  ·  ● deterministic — pure code, testable, no model  ·  ● infrastructure

AI

Voice → structured test data

Whisper transcribes; an LLM extracts test type, readings, pass/fail, weather, and crew into typed fields. On-demand — nothing auto-applies without review.

AI

Voice-first daily report

A 90-second recording plus today’s test entries becomes an AI-generated professional narrative — one report per project per day.

AI

Photo → readings

Photograph a gauge; a vision model extracts the readings straight into the form.

AI

Self-improving transcription

Engineer corrections are logged into a learned lexicon that biases future Whisper runs — a twice-corrected term stops garbling.

Deterministic

Auto-drafted NCRs

A failed test deterministically drafts a linked non-conformance report from a template — no AI — and never overwrites human edits on re-save.

Deterministic

Word template engine

Fills the customer’s own .docx report template in place — byte-identical fonts and letterhead. The app proposes field mappings once; a human confirms.

Deterministic

Field plan daily status

Annotate the site plan PDF; shape colors auto-age by date, replacing manual daily re-coloring in Adobe. Exports as a stamped PDF with legend.

Deterministic

Month-end package

Every daily report, photo, test table, and plan export combined into one client-ready PDF — fully deterministic.

Infrastructure

Offline mode

Full capture with no signal; an IndexedDB queue syncs with retry and idempotency keys when the network returns.

Infrastructure

Two-way Google Docs sync

Reports mirror to Drive every 5 minutes; a human’s edit in the Doc flows back to the database, with zone ownership so data blocks can’t be corrupted.

Infrastructure

Business platform

Issues tracking, monthly AI summaries, Stripe billing, company workspaces with owner dashboards, platform admin, audit logging, per-user feature gates.

03 — Architecture

Serverless app, Postgres as the engine room

Next.js 14 on Vercel; Supabase provides Postgres (with row-level security), auth, file storage, and scheduling. AI is a swappable provider layer — Claude by default, with a self-hosted open-model path as a cost/privacy option. Everything is event-driven or on-demand except one cron-drained sync queue.

System flow — drawn at the boxes-and-arrows level

  1. 1

    Capture in the PWA

    Engineer records voice or snaps a photo → audio hits a transcription route (Whisper, prompt-biased by the learned lexicon).

  2. 2

    Structured extraction

    Transcript → a small fast LLM (Claude Haiku) returns structured JSON → engineer reviews; deterministic heuristics flag suspect words → confirm.

  3. 3

    Guarded write path

    Saves pass through API routes enforcing auth, project ownership, Postgres-backed rate limits, and field allowlists — with row-level security as a second, database-level enforcement layer.

  4. 4

    Triggers react to writes

    Failed test → NCR draft. Report write → sync job enqueued. Business logic lives next to the data.

  5. 5

    Cron-drained sync queue

    pg_cron fires every 5 minutes → a worker claims jobs with FOR UPDATE SKIP LOCKED → renders to Google Docs and pulls human edits back.

  6. 6

    Document generation

    A larger LLM (Claude Sonnet) writes narratives; deterministic code fills Word templates and assembles PDFs. Outputs land in private storage, served via short-lived signed URLs.

  7. 7

    Offline replay

    Writes queue in IndexedDB → on reconnect, replay through the same API routes (so server-side triggers still fire) with client-generated idempotency keys.

04 — Technical Decisions

Choices, alternatives, and why

Two-tier model strategy

Extraction runs on Claude Haiku with tight max_tokens caps (~300–800); narrative generation runs on Sonnet. The alternative — one big model everywhere — pays premium prices for a JSON-shaping task.

Why — the small model is ~5× cheaper and faster at extraction; quality only pays for itself in prose.

No AI where determinism works

NCR drafting, PDF assembly, and transcript-suspect detection (a curated confusion map: “exclamation” → “excavation”) are pure code.

Why — a wrong auto-drafted, client-visible document is a trust cost no accuracy rate justifies. Deterministic code is free and testable.

Fill the customer’s Word template instead of generating PDFs

Engineers already own approved templates. The engine parses the .docx — a hand-rolled ZIP reader/writer, cross-validated against Python’s zipfile, because a malformed zip means “Word cannot open this file” on a client deliverable — and substitutes values in place. Detection is deterministic; classification is human-confirmed once per template.

Why — in the first real template, a project number and a weather value carried identical markup; a model guessing wrong would blank a field on every future report.

RLS as a real enforcement layer, not a backstop

Every mutation is guarded twice — route-level ownership checks and restrictive database policies — because the browser holds a database key by design.

Why — see the hardest problem below. A guard that lives only in application code is not a guard.

Postgres for infrastructure primitives

Rate limiting, job queues, cron, and idempotency all live in Postgres — atomic upserts, SKIP LOCKED, pg_cron — instead of Redis or queue services. Serverless cold starts had already killed in-memory rate limiting.

Why — one database means one thing to reason about transactionally.

Provider-swappable AI with schema enforcement

Local open models don’t follow “return these exact field names” from prose the way Claude does, so the local path passes explicit JSON Schema constraints. Capped tokens, review-before-save, per-user kill-switches, and rate limits on every AI route.

Why — the same product must run on hosted or self-hosted models without a rewrite.

05 — Hardest Problem

The security hole a green test suite couldn’t see

The scariest bug shipped past a fully green test suite. Every API route correctly verified project ownership before inserting — but the routes aren’t the only door. A Supabase app ships its anon database key to every browser by design, and several components wrote to the database directly. The row-level security policies all said “this row is mine” — none said “this project is mine.”

So any signed-in user could insert a test entry carrying their own user ID but someone else’s project ID, and Postgres accepted it: cross-tenant data injection into another company’s reports.

It surfaced only because the verification suite was rebuilt to test as real signed-in users rather than with the service key — which bypasses RLS, so a green service-role run proves nothing about policies. A second trap during diagnosis: probe rows rejected by CHECK constraints looked identical to policy rejections, so refusals had to be asserted by exact error code, not “it failed.” The fix was a database-level restrictive policy — an ownership check ANDed onto every write policy across eleven tables — verified live with two real accounts attacking each other.

A guard that lives only in application code is not a guard. Test at the layer that actually enforces.

06 — Engineering Metrics

Measured, not marketed

1–2h → 90s
daily documentation reduced to one voice recording plus review (est.)
<$0.01
extraction cost per voice log; full AI daily narrative ~$0.02–0.05 (est.)
60
check end-to-end security journey — two real users + a stranger, asserting both “can do” and “must-not-do” — passing
11 / 40
tables with dual-layer write protection, across 40 RLS-enforced migrations
≤5min
save-to-Google-Doc sync latency, with job coalescing and content-hash skips
~19
hand-written E2E verification scripts against the real database and live routes

Solo-built, pre-launch. Subscription pricing (Stripe live), positioned far under Procore’s $375/user/month. Every AI route carries its own hourly per-user rate limit and kill-switch; offline capture runs on a 150MB client-side budget with zero-duplicate replays via idempotency keys.

07 — Tech Stack

What it’s built with

Frontend

  • Next.js 14 (App Router) · React · TypeScript
  • Tailwind + shadcn/ui
  • PWA · IndexedDB offline queue · pdf.js

Backend

  • Next.js API routes (serverless)
  • Supabase Postgres + RLS
  • Postgres triggers · functions · pg_cron

AI / LLM

  • Claude — Haiku extraction, Sonnet generation
  • OpenAI Whisper STT
  • Optional self-hosted: Ollama (Qwen 2.5 72B + vision), whisper.cpp

Documents & Integrations

  • pdf-lib · hand-rolled .docx/ZIP engine
  • Google Docs & Drive APIs (OAuth)
  • Stripe (idempotent webhooks) · Resend

Infra & Quality

  • Vercel · Supabase (auth, signed-URL storage, scheduling)
  • Cloudflare Tunnel for the local-model path
  • ~19 E2E scripts · audit logging · automated spec-review agent gating every change
08 — My Role

Solo, end-to-end

I designed and built everything solo: product scoping against incumbent tools, the data model and every migration, all API routes and security architecture, the AI layer — model selection, prompt design, structured extraction, the transcription-correction learning loop — the offline sync engine, the document engines (Word template fill, PDF generation, Google Docs two-way sync), billing, admin tooling, and deployment.

I also built the QA methodology: an end-to-end verification suite that tests as real users against the live database, and an automated spec-review agent that audits every code change before it ships.

09 — What’s Next

Roadmap

Local model production pilot

Move extraction fully onto self-hosted open models for near-zero marginal AI cost, using the schema-constrained output work already in place.

Close the trust loop on transcription

The learned lexicon already reduces repeat mishearings; next is measuring correction rates per term to quantify it.

Lesson applied forward

Test at the layer that actually enforces — a passing suite that runs with admin privileges proves nothing about security. Every new table now ships with both-directions policy tests from day one.

Next case study
EasyTab AI →

Want the full walkthrough? I’ll demo it live.

Get in touch