FieldLog AI
Speak for 90 seconds at the end of the day — the paperwork is done.
An AI-powered field documentation platform for construction field engineers and QA inspectors. Speech becomes structured test data, a day’s activity becomes a professional report, and a failed test auto-drafts its own non-conformance report. Incumbents like Procore ($375/user/month) are digital clipboards; FieldLog uses AI as the backbone.
Built solo and pre-launch, tested on seeded demo firms. Figures come from code and test runs, not projections.
1–2 hours of paperwork after a 10-hour shift
A field engineer ends every day with hours of manual documentation: transcribing density gauge readings from handwritten notes, writing a narrative daily report, filling the client’s Word template by hand, drafting a non-conformance report for every failed test, and re-coloring yesterday’s markups on the site plan PDF in Adobe. This happens after the shift, often on a phone, sometimes with no signal.
At typical billing rates that’s $50–100/day of unbillable time per engineer (est.) — and the error mode is worse than the cost: a reading mis-transcribed into a stamped engineering document. Existing tools digitize the clipboard but still make the human do the writing.
AI where it helps, deterministic code where it must be right
● AI — model in the loop, always human-reviewed · ● deterministic — pure code, testable, no model · ● infrastructure
Voice → structured test data
Whisper transcribes; an LLM extracts test type, readings, pass/fail, weather, and crew into typed fields. On-demand — nothing auto-applies without review.
Voice-first daily report
A 90-second recording plus today’s test entries becomes an AI-generated professional narrative — one report per project per day.
Photo → readings
Photograph a gauge; a vision model extracts the readings straight into the form.
Self-improving transcription
Engineer corrections are logged into a learned lexicon that biases future Whisper runs — a twice-corrected term stops garbling.
Auto-drafted NCRs
A failed test deterministically drafts a linked non-conformance report from a template — no AI — and never overwrites human edits on re-save.
Word template engine
Fills the customer’s own .docx report template in place — byte-identical fonts and letterhead. The app proposes field mappings once; a human confirms.
Field plan daily status
Annotate the site plan PDF; shape colors auto-age by date, replacing manual daily re-coloring in Adobe. Exports as a stamped PDF with legend.
Month-end package
Every daily report, photo, test table, and plan export combined into one client-ready PDF — fully deterministic.
Offline mode
Full capture with no signal; an IndexedDB queue syncs with retry and idempotency keys when the network returns.
Two-way Google Docs sync
Reports mirror to Drive every 5 minutes; a human’s edit in the Doc flows back to the database, with zone ownership so data blocks can’t be corrupted.
Business platform
Issues tracking, monthly AI summaries, Stripe billing, company workspaces with owner dashboards, platform admin, audit logging, per-user feature gates.
Serverless app, Postgres as the engine room
Next.js 14 on Vercel; Supabase provides Postgres (with row-level security), auth, file storage, and scheduling. AI is a swappable provider layer — Claude by default, with a self-hosted open-model path as a cost/privacy option. Everything is event-driven or on-demand except one cron-drained sync queue.
System flow — drawn at the boxes-and-arrows level
-
1
Capture in the PWA
Engineer records voice or snaps a photo → audio hits a transcription route (Whisper, prompt-biased by the learned lexicon).
-
2
Structured extraction
Transcript → a small fast LLM (Claude Haiku) returns structured JSON → engineer reviews; deterministic heuristics flag suspect words → confirm.
-
3
Guarded write path
Saves pass through API routes enforcing auth, project ownership, Postgres-backed rate limits, and field allowlists — with row-level security as a second, database-level enforcement layer.
-
4
Triggers react to writes
Failed test → NCR draft. Report write → sync job enqueued. Business logic lives next to the data.
-
5
Cron-drained sync queue
pg_cronfires every 5 minutes → a worker claims jobs withFOR UPDATE SKIP LOCKED→ renders to Google Docs and pulls human edits back. -
6
Document generation
A larger LLM (Claude Sonnet) writes narratives; deterministic code fills Word templates and assembles PDFs. Outputs land in private storage, served via short-lived signed URLs.
-
7
Offline replay
Writes queue in IndexedDB → on reconnect, replay through the same API routes (so server-side triggers still fire) with client-generated idempotency keys.
Choices, alternatives, and why
Two-tier model strategy
Extraction runs on Claude Haiku with tight max_tokens caps (~300–800); narrative generation runs on Sonnet. The alternative — one big model everywhere — pays premium prices for a JSON-shaping task.
Why — the small model is ~5× cheaper and faster at extraction; quality only pays for itself in prose.
No AI where determinism works
NCR drafting, PDF assembly, and transcript-suspect detection (a curated confusion map: “exclamation” → “excavation”) are pure code.
Why — a wrong auto-drafted, client-visible document is a trust cost no accuracy rate justifies. Deterministic code is free and testable.
Fill the customer’s Word template instead of generating PDFs
Engineers already own approved templates. The engine parses the .docx — a hand-rolled ZIP reader/writer, cross-validated against Python’s zipfile, because a malformed zip means “Word cannot open this file” on a client deliverable — and substitutes values in place. Detection is deterministic; classification is human-confirmed once per template.
Why — in the first real template, a project number and a weather value carried identical markup; a model guessing wrong would blank a field on every future report.
RLS as a real enforcement layer, not a backstop
Every mutation is guarded twice — route-level ownership checks and restrictive database policies — because the browser holds a database key by design.
Why — see the hardest problem below. A guard that lives only in application code is not a guard.
Postgres for infrastructure primitives
Rate limiting, job queues, cron, and idempotency all live in Postgres — atomic upserts, SKIP LOCKED, pg_cron — instead of Redis or queue services. Serverless cold starts had already killed in-memory rate limiting.
Why — one database means one thing to reason about transactionally.
Provider-swappable AI with schema enforcement
Local open models don’t follow “return these exact field names” from prose the way Claude does, so the local path passes explicit JSON Schema constraints. Capped tokens, review-before-save, per-user kill-switches, and rate limits on every AI route.
Why — the same product must run on hosted or self-hosted models without a rewrite.
The security hole a green test suite couldn’t see
The scariest bug shipped past a fully green test suite. Every API route correctly verified project ownership before inserting — but the routes aren’t the only door. A Supabase app ships its anon database key to every browser by design, and several components wrote to the database directly. The row-level security policies all said “this row is mine” — none said “this project is mine.”
So any signed-in user could insert a test entry carrying their own user ID but someone else’s project ID, and Postgres accepted it: cross-tenant data injection into another company’s reports.
It surfaced only because the verification suite was rebuilt to test as real signed-in users rather than with the service key — which bypasses RLS, so a green service-role run proves nothing about policies. A second trap during diagnosis: probe rows rejected by CHECK constraints looked identical to policy rejections, so refusals had to be asserted by exact error code, not “it failed.” The fix was a database-level restrictive policy — an ownership check ANDed onto every write policy across eleven tables — verified live with two real accounts attacking each other.
A guard that lives only in application code is not a guard. Test at the layer that actually enforces.
Measured, not marketed
Solo-built, pre-launch. Subscription pricing (Stripe live), positioned far under Procore’s $375/user/month. Every AI route carries its own hourly per-user rate limit and kill-switch; offline capture runs on a 150MB client-side budget with zero-duplicate replays via idempotency keys.
What it’s built with
Frontend
- Next.js 14 (App Router) · React · TypeScript
- Tailwind + shadcn/ui
- PWA · IndexedDB offline queue · pdf.js
Backend
- Next.js API routes (serverless)
- Supabase Postgres + RLS
- Postgres triggers · functions · pg_cron
AI / LLM
- Claude — Haiku extraction, Sonnet generation
- OpenAI Whisper STT
- Optional self-hosted: Ollama (Qwen 2.5 72B + vision), whisper.cpp
Documents & Integrations
- pdf-lib · hand-rolled .docx/ZIP engine
- Google Docs & Drive APIs (OAuth)
- Stripe (idempotent webhooks) · Resend
Infra & Quality
- Vercel · Supabase (auth, signed-URL storage, scheduling)
- Cloudflare Tunnel for the local-model path
- ~19 E2E scripts · audit logging · automated spec-review agent gating every change
Solo, end-to-end
I designed and built everything solo: product scoping against incumbent tools, the data model and every migration, all API routes and security architecture, the AI layer — model selection, prompt design, structured extraction, the transcription-correction learning loop — the offline sync engine, the document engines (Word template fill, PDF generation, Google Docs two-way sync), billing, admin tooling, and deployment.
I also built the QA methodology: an end-to-end verification suite that tests as real users against the live database, and an automated spec-review agent that audits every code change before it ships.