Singaporebuildsai Start a brief

Independent AI practice · one-north, Singapore

Build gallery

Twelve composite patterns for the kinds of systems I actually build. Each one is a shape: the task, what gets in the way, how the path is put together, and how it is checked.

These are composite patterns drawn from the kinds of work I do. They are illustrations of shape and method, not client reports.

Retrieval & search

Policy answer desk

A policy desk answered staff questions by searching a shared drive and hoping the filename was honest. Answers went out without a page reference, so a second person had to re-read the source.

The build indexes the current policy pack, retrieves passages, and requires a citation before an answer can be copied. The desk can open the page and check the wording before sending.

Checks run against a set of questions the desk already knows are awkward, including superseded policies that must not be quoted.

retrieval over policy PDFscitation requiredstale-policy trap cases

Retrieval & search

Tender clause finder

Bid teams spent evenings hunting for a clause they remembered from an earlier pack. Keyword search failed when the wording had been negotiated into something adjacent.

The system retrieves candidate clauses from past tenders and from the current draft, then lines them up so a reviewer can accept, edit, or reject. Nothing is pasted into a bid without that step.

The harness includes clauses that look similar and are legally different, which is where naive retrieval is most confident and most wrong.

clause alignmenthuman accept/edit/rejectnear-duplicate traps

Retrieval & search

Field manual assistant

Technicians in the field needed a procedure from a manual that lives as a mix of PDFs and scanned pages. Phone search returned the wrong revision more often than anyone wanted to admit.

The assistant retrieves from a pinned revision, shows the figure or step, and labels the revision on every card. If the index is behind the published pack, the card says so.

Evaluation uses the questions supervisors already ask on a Monday briefing, plus a handful of procedures that were rewritten in the last quarter.

pinned revisionfigure + step retrievalrevision label on every card

Document automation

Invoice-to-ledger pipeline

Supplier PDFs arrived in a mailbox and were typed into the ledger by hand. Totals sometimes disagreed with the purchase order, and the disagreement was noticed late.

The pipeline extracts fields, checks them against a vendor list and an open PO, and parks the file when the totals do not match. A person handles the parked queue; the model never posts.

The test set is a folder of invoices that already caused trouble, including stamped scans and split line items.

field extractionPO checkhuman exception queue

Document automation

KYC pack pre-check

Onboarding packs arrived incomplete, and the gap was found after a reviewer had already started. Missing pages and expired IDs were the usual reasons.

A pre-check lists what is present, what is unreadable, and what has passed a stated expiry. It does not decide whether a customer is acceptable. That decision stays with the reviewer.

We test on packs that were previously sent back, so the harness fails if the pre-check goes quiet on a missing document.

presence checklistexpiry flagno eligibility decision

Document automation

Shipping document parser

Bills of lading, packing lists, and certificates arrived as a zip of mixed scans. Ops re-keyed vessel names and container numbers, then reconciled them at the gate.

The parser pulls the identifiers, matches them across the pack, and highlights a mismatch before the goods move. Unreadable stamps go to a person, with the original page attached.

Nightly regression uses packs that already produced a mismatch in the last year, including rotated scans.

cross-document matchunreadable-stamp pathnightly regression run

Forecasting & ops

Spare-parts demand window

Parts were ordered by habit and by whoever shouted first. Stock-outs and over-orders arrived in the same month, and the spreadsheet that tried to keep up had one owner.

A weekly window gives a suggested band for each SKU, with the error shown in units the store already counts. The planner can override; the override is logged and becomes next month’s training case.

We hold out a quiet quarter and a messy quarter, so the model is judged on both kinds of week.

weekly banderror in operational unitsoverride log

Forecasting & ops

Service queue triage

A service desk tagged tickets by reading the first two sentences. Urgent work hid under polite language, and a few categories had become a dumping ground.

The model suggests a queue and a urgency band, with the original text still on screen. A lead can pin a rule that always wins: some senders, some words, some contract IDs.

Evaluation uses tickets the leads already marked as mis-routed, which is a harsher set than a random sample.

suggested queuepin-a-rule overridemis-route test set

Forecasting & ops

Rostering pressure forecast

A roster lead could feel a bad week coming and still lacked a number to take into the planning meeting. Leave, seasonality, and a few large clients moved the load together.

The forecast shows pressure by desk for the next fortnight, with the drivers listed in plain language. It does not write the roster. It gives the meeting something to argue with.

We score it on weeks the lead already remembers as painful, including public-holiday stretches.

fortnight windowlisted driversno auto-roster

Evaluation & guardrails

Regression suite for a support bot

A support bot had been tuned in chat until nobody could say whether last week’s prompt change had helped. Complaints arrived anecdotally.

A pytest suite now runs a fixed set of threads: refunds, account locks, and the questions the bot must refuse. A change that drops a case is a failed build, even if the demo looks smoother.

The suite is the handover. The bot is allowed to change only while the suite stays green.

pytest eval runnermust-refuse caseschange gated on green

Evaluation & guardrails

Grounding and citation checks

An internal assistant sounded fluent and still invented a paragraph that was not in the corpus. Reviewers caught it late, after a message had already been forwarded.

A check now fails the answer if a claim cannot be aligned to a retrieved passage. Uncited sentences are stripped or sent back. A nightly job repeats the check on the golden set.

Answers now carry a page reference, so the desk can verify before sending.

passage alignmentcitation checknightly regression run

Internal copilots

Sales desk briefing copilot

Account managers opened four systems before a call and still missed a renewal date. The briefing, when it existed, was a paste of the last email.

The copilot pulls the account record, the open tickets, and the last meeting note into a one-page brief with sources. It does not draft the commercial offer. That stays with the manager.

We score briefs on whether the renewal date, the open risk, and the last commitment are present, which is what the desk said it actually needed.

one-page briefsource linksno offer drafting

What a build includes

  1. A repository with a README. The why of each choice sits next to the how. A stranger to the project should be able to run the evals on a working day.
  2. An environment. Dev and a production-shaped path, with secrets kept in your store. I do not leave credentials in the repo.
  3. A test set from your real examples. We pick cases together, including the ugly ones. This set is versioned. It is the definition of done.
  4. A runbook. How to deploy, how to replay a failure, who to call when the index is stale. Written for the two people who will own it.
  5. A cost log. Tokens, hosts, and scheduled jobs, next to the quality line, so scale is a decision rather than a surprise.
  6. Training for two of your people. A live session on the repo, the harness, and the exception path. Recorded if you want it.
  7. Thirty days of support after transfer. Defects in what I shipped are mine to fix. New scope is a new conversation.

If one of these shapes is close

Tell me which desk and which documents. I will say whether the pattern still holds once I have seen a sample of the data.

Start a brief