Singaporebuildsai Start a brief

Independent AI practice · one-north, Singapore

Notes from the build desk

Short pieces from 2026 about how I actually do the work. They are written from the desk in one-north, for people who own a workflow and a test set.

January 2026 · Practice

Start with the evaluation set, not the model

The first meeting on a new brief often arrives with a model already chosen. Someone saw a demo. Someone else has a licence sitting unused. The workflow is described in adjectives. When I ask for twenty examples of a good output, the room goes quiet, and that quiet is useful. It tells me we are still in scoping, which is a cheaper place to be honest.

An evaluation set is a pile of real inputs and the outputs a named person would accept. The awkward cases belong in it: the superseded policy, the unreadable stamp, the ticket written in three languages. If those cases are left for “later”, the prototype will look strong and the desk will not trust it. I version the set with the code, so a prompt change that drops a case is a failed build rather than a matter of taste.

Writing the set also names the owner. Someone has to say whether an answer is acceptable in the actual job. If that person cannot spare the time, we should not start. I would rather pause for a month than invent a gold standard from a vendor card.

Once the set exists, picking a model is a smaller decision. We run the smallest thing that passes, then stop. The set is allowed to grow as the desk finds new failures. That growth is the work. The model is a detail we can swap when the harness says we may.

I now refuse to quote a quality number until this artefact exists. Teams that wanted a model first sometimes leave after that conversation. Teams that stay tend to finish.

Back to the top of the notes

February 2026 · Evaluation

What retrieval actually fails at

When an internal assistant gives a wrong answer, the instinct is to change the generator. I have stopped starting there. The last three broken systems I was asked to look at failed for reasons that had little to do with the model at the end of the pipe. The files were chunked on page breaks that split a clause in half. The index still held a policy that legal had withdrawn. A share permission meant the retriever could not see the document the desk was allowed to read.

Those failures look like hallucinations from the outside. They are closer to a filing problem. A citation check will catch some of them, if the cited passage is the wrong revision. It will not catch a missing file. The harness has to include questions whose answers live in documents that were updated last week, and questions whose answers used to live in documents that must no longer be quoted.

Chunking deserves more of the week than it usually gets. A tender clause, a procedure with a figure, and a table of rates each want a different split. I keep a small set of “near-duplicate” questions in the evals, where two passages look similar and only one is valid. Naive retrieval is most confident on those, and most wrong.

Only when chunking, permissions, and freshness are boringly correct do I look at the generator. Sometimes a smaller model then passes. That is a cheaper finding than another fortnight of prompt theatre.

Back to the top of the notes

February 2026 · Data

Keeping client data inside Singapore

Teams here often have a residency requirement they can state in one sentence: these records should not leave Singapore. I treat that sentence as an architecture input. It decides whether a hosted model API is even on the table, where the index lives, and whether the evaluation runner may call out. I design around that constraint in the stack.

Singapore’s Personal Data Protection Act sits in the background of that conversation. I am not a lawyer and I am not your data protection officer. I do not certify compliance, and this note is not legal advice. What I can do is ask where personal data is present in the corpus, who is allowed to see it, and which regions are forbidden. Those answers change the build.

A practical pattern is to run retrieval and the model in an environment you already control, in-country, and to keep the test set there as well. Hosted APIs are used only when you have already accepted that path in writing. Logs need the same care as the documents; a helpful trace that stores a passport page is a new copy of the thing we were trying to contain.

The harness has to live under the same rule. If the nightly job ships examples to a region you have not approved, the evaluation itself becomes the leak. I now write the residency assumption at the top of the README, next to the model choice, so the person who inherits the repo in a year does not “just try” a foreign API to see if quality improves.

If you need a legal opinion, you need counsel. If you need a system that respects the restriction you have already been given, that is work I know how to do.

Back to the top of the notes

March 2026 · Practice

The cost line nobody models

A prototype on twenty examples is almost always cheap. The bill appears when the desk is busy, the index is rebuilt overnight, and every answer is re-checked for a citation. I have watched a path that felt free in week two become a line item in week eight, after nobody had written the cost down next to the quality score.

I keep a cost log from the first harness run. Tokens, host time, scheduled jobs, and the human minutes in the exception queue. The last of those is the one teams forget. A pipeline that parks every other invoice for a person has not automated the desk. It has moved the typing into a folder with a worse user interface.

Smaller models help, which is one reason I start there. Caching retrieval results helps. Refusing to generate when the retriever is empty helps more than a clever prompt. Each of those choices is visible on the cost line. If the line is only in my notebook, you cannot decide whether to scale.

The handover includes that log and a way to keep it. A quality number without a cost number is half a measurement. I will not recommend a production cut-over until both are in the weekly note you already know how to read.

Back to the top of the notes

April 2026 · Tooling

Small models are often enough

Once a harness exists, the interesting question is the smallest model that still passes it. I have shipped retrieval assistants and parsers on small open-weight models running in an environment the client already operated. They were slower to write a sentence and faster to approve, because the residency story was simple and the cost line stayed still.

Large hosted models are a tool I will use when the set demands them and when you have permitted the data path. They are a poor default. They hide the fact that the retrieval is wrong, because the generator is good at sounding like the missing passage. They also make the nightly eval more expensive, which means people run it less, which means quality drifts without a witness.

The promotion rule in my repos is dull: a larger model is allowed when the harness says the smaller one cannot pass, and when we have already tried chunking, prompts, and the obvious rules. Dull rules keep the README honest. A year later, someone will want to swap the model because a new card was published. They can, if the suite stays green.

I still keep a hosted API in the toolkit. Some tasks, on some corpora, need it. The point is that “need” is a measurement, and the measurement lives in pytest.

Back to the top of the notes

May 2026 · Practice

Handover is the deliverable

A system that only runs on my laptop is a conversation. The product is the moment two of your people can deploy, replay a failure, and read a red eval without calling me. Until that moment, I have not finished, however tidy the demo looked on Friday.

Handover, as I write it into a statement of work, has four parts. The repository in your account, with a README that explains the choices I rejected. An environment you control. The test set drawn from your examples. A runbook and a live session with two named people. Thirty days of support follows, for defects in what I shipped. New scope is a new quote.

I ask for those two names in week one. If they cannot be found, we should stay in discovery. A build with no operator becomes a folder that nobody opens, which is a more expensive way to keep the original workflow.

The session is practical. We break the eval on purpose, restore it, and walk the exception queue. I want the discomfort in the room while I am still there. After that, the practice’s job is to get out of the way. You should be able to change a prompt, fail a test, and fix it with the notes already in the repo.

That is also why I take one engagement at a time. Handover needs attention. Splitting that attention across two desks is how demos ship and runbooks do not.

Back to the top of the notes