I want each product I build to make the next one easier to build well. That is why I am building a product factory. For the kind of work I want to do, it is the best approach: make the product decisions explicit, give agents bounded work, and retain evidence that the result actually works.
StillRoom put that approach under pressure. The idea was a native Mac app for screenshots scattered across folders. The outcome was simple: find and reuse a capture without remembering its filename or where it was saved. Getting there involved local machine learning, file safety, an unsuccessful first attempt at visual similarity, and a release process that the factory's existing web pipeline could not provide.
Start with the thing a person needs to do
StillRoom's native source definition names the user and the outcome before listing capabilities. Here is an excerpt from the actual spec:
{
"product_id": "stillroom",
"platform": "macos",
"minimum_os": "13.0",
"brief": {
"name": "Stillroom",
"primary_user": "Mac users with screenshots scattered across folders",
"outcome": "Find and reuse a saved screenshot without remembering its filename or location"
}
}That outcome narrowed the product. Choose folders explicitly. Keep originals where they are. Search recognized text, filenames, and personal notes. Select an image to look for related captures. Open the original, copy its text or image, and return to the work that made the screenshot useful.
The implementation uses SwiftUI and AppKit, with Apple Vision and Natural Language for local analysis. It has no account requirement, screenshot upload path, server, or telemetry. OCR and ordinary text search remain useful when the local sentence-embedding model is unavailable. The product does not need a chatbot to do its job.
Write down what must stay true
A feature list leaves too many decisions to the implementation. "Find duplicates" does not say whether similar-looking files count, whether originals can be deleted, or what happens if a file changes after indexing.
The product brief answers those questions with acceptance criteria. This shortened rendering preserves the substance of several actual requirements; the YAML is an explanatory format, not a separate executable factory schema:
SR-01:
scope: Enumerate only explicitly selected folders.
boundary: Exclude symlinks.
SR-02:
fallback: Keep OCR search when embeddings are unavailable.
SR-04:
exact_copy: Equal SHA-256 hashes at different paths.
trash: Confirm one file, then check its current bytes again.
SR-05:
persistence: Save the index and user notes atomically.
failure: Never replace an unreadable index with an empty one.These rules shape architecture. StillroomCore owns indexing, retrieval, persistence, and safe file actions. The native app owns presentation and explicit user actions. StillroomLab generates synthetic captures and exercises the composed system. A decision about deleting a file belongs with the file-safety rules, where it can be tested independently of the screen containing the button.
This is a factory principle I want to keep: specify the valuable outcome and the invariants together. Agents then have a concrete definition of success, including the mistakes they must prevent.
The useful lesson came from a bad result
The first visual-similarity implementation returned the nearest 24 available feature prints. It could always fill a grid. An image with no useful neighbors still received suggestions, because ranking the available candidates did not establish that any candidate belonged in the results.
The first repair imposed a conservative distance cutoff. That removed unrelated suggestions, but it also hid a coherent group of related captures. The evidence showed why one cutoff was insufficient: some relevant pairs were farther apart in feature space than an irrelevant pair.
The next rule combined visual distance with supporting context. A wider cutoff requires the same full folder path, captures within 18 hours, and at least two shared image labels beyond generic terms such as "screenshot" or "document." Exact byte copies qualify separately. The result limit applies after admission, and an empty result is allowed.
A retained audit reported eight related captures for the positive example and zero for the unrelated example. That is evidence about those examples, not a universal retrieval benchmark. The contextual thresholds remain prototype judgments. What matters for the factory is the loop: retain the failure, revise the rule, test the distinction, and make the result understandable in the interface.
Make tests earn their place
The September 25 validation record reports 14 passing tests, divided into nine unit tests and five component tests. The categories matter because a pure retrieval policy and real filesystem indexing prove different things. Tests cover exact-text priority, embedding fallback, corrupt-index protection, changed-file safety, and preservation of notes across rescans.
The work also challenged tests with deliberate faults. Removing lexical priority, weakening duplicate identity, skipping the changed-byte check, and omitting note preservation each caused the relevant behavioral assertion to fail. The retained mutation report records nine detected faults. Those faults were restored before the passing build.
The local journey used eight synthetic screenshots through real OCR, classification, embeddings, visual features, persistence, and repeat scanning. OCR found ECONNREFUSED; an exact-copy group was retained; source bytes stayed unchanged. A search for "breakfast ingredients" found the recipe as a semantic suggestion. A travel-related query returned nothing. Recording both outcomes keeps the evidence useful.
Let each product improve the factory
StillRoom is native software. The factory's hosted create-product pipeline did not produce this app. The work added a native source definition and package-verification path instead of treating a successful web deployment as evidence that a Mac application was ready.
The native package verifier runs tests, builds the app, checks bundle identity and signature, and retains the bundle with a verification report. Distribution verification is a separate step. The commands make that separation visible:
PYTHONPATH=src python3 -m product_factory_runtime verify-native-package \
--spec specs/stillroom.native.json
PYTHONPATH=src python3 -m product_factory_runtime verify-native-distribution \
--spec specs/stillroom.native.json \
--package-report PATH_TO_VERIFIED_PACKAGE_REPORTThe important distinction is captured in the actual version record:
{
"version": "0.1.0",
"source_status": "locally_verified",
"distribution_status": "not_published",
"factory_product_intake_status": "pending_authorized_intake"
}A local app bundle does not prove a public release. StillRoom still needs Developer ID signing, notarization, and a fresh-Mac installation check before public distribution. Large-library performance, broader relevance evaluation, and a complete accessibility audit remain open. Those limits belong in the case study because they determine what another person can safely expect from the product.
Why I choose the factory approach
The payoff is that product knowledge survives the coding session. The next native product can reuse package verification and release evidence. The next retrieval feature can reuse the lesson that ranking needs an admission rule. The next agent can start from a named behavior, an owning component, and a testable invariant.
My factory principles are straightforward: start from a customer outcome; keep product-specific decisions in the product; give components clear ownership; test observable behavior; preserve evidence of failures as well as successes; and distinguish local validation from distribution readiness. Reuse the engineering process while allowing each product to have its own shape.
That is why I see a product factory as the best way to build a portfolio of products with agents. Each build can improve both the product and the process used to build it. StillRoom is an early, concrete example of that approach: a useful local beta, a corrected retrieval rule, and a native verification capability the factory can use again.
Source note: the spec excerpts and reported results come from StillRoom 0.1.0's native source definition, product brief, version record, validation report, and similarity-repair notes. Validation results are reported from retained runs, not a new benchmark conducted for this article. The product screenshots show synthetic sample captures. No personal screenshots or personal file paths are included.