---
name: consumer-ai-demand-probe-prototype
description: "Build a fast Consumer AI Factory demand-probe prototype from an opportunity scan: choose the highest-signal wedge, implement only the first-value loop, verify with tests/build/browser QA, redesign if feedback says the UI lacks credibility, and log artifacts in Obsidian."
version: 1.0.0
author: Hermes Agent
metadata:
  hermes:
    tags: [consumer-ai-factory, prototype, demand-probe, product-strategy, frontend]
    related_skills: [consumer-ai-factory-opportunity-evaluation, test-driven-development, popular-web-designs, obsidian]
---

# Consumer AI Demand-Probe Prototype

## When to use

Use when Antoine asks to “build the prototype,” “make a quick MVP,” “test this app idea,” or move a Consumer AI Factory opportunity from research into a concrete artifact.

This skill is for demand-probe prototypes, not full products.

## Core doctrine

Do not overbuild.

If Antoine says a landing page/prototype is “pretty good,” “good,” or otherwise approves the direction after QA, stop polishing unless he asks for edits. Treat the artifact as ready for a demand test and shift to capture + traffic: verify waitlist/payment/submission storage, install/verify telemetry, connect a credible URL if needed, and send qualified traffic. Metrics and user objections should drive the next design pass, not internal perfectionism.

The right first build is usually not a sophisticated AI agent. It is an interactive artifact that makes the promised first-value output concrete enough to test whether users will:

1. understand the promise,
2. enter/upload real material,
3. feel the output is valuable,
4. show payment intent.

For Consumer AI Factory work, the prototype should answer: “Does the wedge deserve a demand probe?” not “Can we build the whole app?”

### Antoine/Mr Han testing-first posture

Antoine wants Mr Han oriented toward testing, QA, validation evidence, and product truth more than broad venture operation. For prototype/MVP work, default to acting as a skeptical product tester: define what must be proven, build only enough artifact to expose that risk, verify in browser, report what broke trust or blocked activation, and recommend the next evidence-gathering action. Avoid drifting into ongoing venture ops, feature accumulation, or broad growth work unless Antoine explicitly asks for operation rather than testing.

When Antoine asks a quick status or ideation question, lead with the answer first; do not run a full audit/build loop unless the wording clearly asks for verification, fixing, deployment, or production QA.

## Workflow

1. Pick the highest-signal wedge
   - Read the relevant opportunity scan / research artifact.
   - Choose the opportunity with the clearest first-use moment, not necessarily the cleverest idea.
   - For the 2026-05-10 scan, Homeowner Triage was chosen because the first-value loop was clean: messy inspection/report input -> prioritized safety-first 90-day plan.

2. Define the prototype promise
   - One sentence: input -> output.
   - Example: “Paste inspection notes or repair quote text -> get emergency/this-week/this-month/DIY priorities, a now/30/90-day plan, contractor questions, and quote sanity checks.”
   - If the prototype risks being “just ChatGPT with a wrapper,” stop and sharpen it into persistent workflow/decision support before building more UI. For Homeowner Triage, the report-triage wedge remained low-conviction because a motivated user could paste the same report into ChatGPT and get a decent answer.
   - Do not force “human-reviewed in 24h” as the core offer just because it helps a demand probe. It only makes sense when trust is genuinely the wedge; otherwise it conflicts with user expectations that AI should be instant.
   - Before continuing prototype polish, verify two strategic gates:
     1. Distribution-owner map: who already aggregates this audience with trust?
     2. Non-ChatGPT advantage: what does the product remember, ingest, organize, track, or complete that chat does not?
   - If those gates are weak, park the prototype and reframe the opportunity instead of polishing.

3. Build only the first-value loop
   - Avoid auth, accounts, database, subscriptions, onboarding complexity, dashboards, and “real AI” unless required.
   - A rule-based parser or hardcoded demo can be valid if the goal is testing output shape and upload/payment behavior.
   - Make the output specific enough to feel real.
   - Exception for fully automated visual AI products: if the venture lives or dies on model output quality (e.g. face photo -> realistic haircut preview, body photo -> styling preview), do not start with a landing page or UI shell. Start with a technical feasibility spike that tests the core generation loop directly. Build a batch script/repo that takes representative consented inputs, tries multiple providers/models, saves comparison grids, logs prompts/settings/cost/latency, and scores outputs against a concrete acceptance rubric before any auth/payments/landing-page polish.
   - For identity-preserving photo products, the first spike should usually compare hosted editors first (e.g. OpenAI/GPT Image edit, FLUX/Kontext, Stability inpaint via fal/Replicate/direct APIs), then add segmentation/masking and identity QA only if raw edits drift. Use face/hair parsing, conservative edit masks, face-similarity checks, and reject/regenerate loops as the defensible product layer over commodity image APIs.

4. Use TDD for core logic
   - Load and follow `test-driven-development` when implementing behavior.
   - Write failing tests before production logic for classification, prioritization, formatting, and empty states.
   - Then implement the minimum logic to pass.
   - Verify with the project test command.

5. Design for credibility, not AI hype
   - If the category involves trust/safety/money/home/health/legal-ish decisions, avoid dark neon/glowy “AI demo” visuals.
   - Prefer restrained, warm, credible service UI.
   - Do not start coding UI from taste, copy, or a blank CSS file. First define the product object, user anxiety, conversion moment, and visual reference set.
   - Do not use random design-reference shorthand as if it is product direction. Antoine does not want vague prompts like “make it Apple Settings + Wise.” Translate references into artifact-level decisions: what the user sees first, what is clickable, what information is visible, what gets removed, what anxiety is handled where, and why this is materially better than the old screen.
   - Load `popular-web-designs` only as calibration, not as a substitute for Antoine’s own good/bad examples. When Antoine provides examples, capture them in the Taste Reference Library and treat them as higher-priority than generic reference templates.
   - Maintain and consult the Consumer AI Factory taste/reference artifacts before major UI work:
     - `/Users/antoinelevy/Documents/Obsidian Vault/Consumer AI Company Factory/01 Playbooks/Taste Reference Library.md`
     - `/Users/antoinelevy/Documents/Obsidian Vault/Consumer AI Company Factory/01 Playbooks/a16z Top 100 Gen AI Apps - UX Reference Scan.md`
   - Use these references as concrete pattern libraries, not mood boards. Extract artifact-level rules such as: hero-as-product, direct trial interaction, finished artifact before feature list, atmosphere plus specificity, input-as-CTA for creation tools, source/uncertainty proof for sensitive tools, and visible transformation for visual AI tools.
   - A strong consumer AI first viewport should usually compress four jobs: emotional promise, direct product interaction, product/output artifact proof, and low-friction CTA. If one of these is missing, call it out before coding.
   - Antoine’s positive examples so far:
     - HappyCouple hero: dark/private/warm palette, central pod, hover/click affordance, direct “click to talk” test, emotionally specific H1, aligned CTAs, real prompt examples.
     - character.ai signup: simple, visually proper, cinematic image + familiar signup card, but information-light; use atmosphere only if paired with product specificity.
     - Suno: hero is the creation surface; prompt composer + create button + polished playable song artifacts.
     - Photoroom: outcome-first commerce promise, immediate visual product proof, visible editing controls, transformation/output examples, layered credibility.
   - Before writing UI code, create a UX/art-direction brief:
     1. target user and core usage moment,
     2. primary product object users manage or receive,
     3. the hero’s direct product interaction — the prototype’s equivalent of the HappyCouple pod: a concrete, clickable object that lets the user test/feel the product immediately,
     4. best-in-class references and anti-references translated into concrete screen rules, not vibe labels,
     5. spacing/type/card/CTA system,
     6. trust objections that must appear near the relevant action,
     7. mobile behavior and first viewport hierarchy,
     8. what must be removed from the previous version so the rebuild is not same-skeleton polish.
   - Use established consumer-app patterns: one dominant product/interaction object, one primary action, explicit status/next step, semantic color, tokenized spacing, visible empty/loading/error states, and source/audit trails for sensitive AI outputs.
   - If the prototype's main interaction is a giant editable paste/chat box, treat it as suspect. Reframe around before/after proof and a saved/usable artifact before exposing raw input controls.
   - Customer-facing copy should not say “demand probe,” “prototype,” or “this is a test” in prominent page sections unless the audience is internal/startup-savvy.
   - Never expose internal payment-intent controls like “I would pay for this” as a public CTA. That is a founder metric, not a user action. Translate it into a real next step the user understands:
     - Bad: “I would pay for this.”
     - Better: “Request human-reviewed plan — $49.”
     - Bad: “Join demand test.”
     - Better: “Get secure upload instructions.”
   - The CTA must answer what happens next: checkout, secure upload, email follow-up, sample output, booking, or download.
   - Avoid “private beta” as the default conversion frame when the real question is willingness to pay. If payment is the desired evidence, use a paid setup/checkout-style offer (e.g. “Start for $19”) and be explicit internally whether real payment is wired. LocalStorage intent capture is only prototype-level; before sending external traffic, wire Stripe/payment + secure upload or clearly treat the page as a local mock.
   - Translate founder-facing language into user-facing benefit:
     - Bad: “Built to test demand.”
     - Better: “Organize before you call.”
     - Bad: “This is the actual test: will users pay?”
     - Better: “For recent buyers who want clarity before making repair decisions.”

6. Run browser QA, not just build checks
   - Start a local preview on an explicit port.
   - If framework dev server is flaky, build static output and serve it with Python `http.server` as a visual preview fallback.
   - Use browser snapshot for interaction checks.
   - Use browser vision for visual QA and ask specifically whether the page looks credible and whether there are glaring layout issues.
   - Treat visual quality as a release gate, not optional polish. Before reporting back, self-reject screens that have obvious UX failures: dashboard/product proof does not fit a normal viewport, confusing step numbers or badges, oversized/fat CTAs, cramped/cropped cards, blank high-friction forms, unclear CTA hierarchy, repeated competing CTAs, or fake/prototype-smelling copy.
   - For consumer prototypes, run at least one harsh desktop visual pass around a normal laptop viewport and verify the product proof + conversion path are understandable without excessive scrolling. If mobile matters for the channel, run a mobile-width pass too.
   - If user says the design is bad, accept it and redesign immediately; do not defend the first pass.
   - If Antoine says one product is “good for now” and asks to apply what was learned to other products, do not inspect or modify the good product. Treat it as the reference pattern only; rebuild the named target artifacts and report only those paths.
   - If Antoine says a redesign “looks the same,” accept that as the relevant UX signal. Do not defend browser QA that only says the page is acceptable. The bar becomes material difference: build a separate route/file (for example `/v2.html`), keep the old version viewable, compare old vs new screenshots side-by-side, and ask whether a normal user would instantly see a different/better product experience.
   - Do not call same-skeleton polish a rebuild. Copy changes, trust copy, CTA labels, slightly better cards, and more source proof are revisions unless the first viewport, product object, and interaction model clearly changed.
   - If local HTTP preview servers crash/refuse connections even after Python/Vite/http-server fallbacks, do not skip visual QA. Add `base: './'` to Vite config if needed, generate a temporary `dist/qa-inline.html` with built CSS/JS inlined, load it via `file://`, and clearly treat it as a QA artifact rather than source of truth.
   - If the browser MCP/session is unavailable during local prototype QA, do not downgrade to build-only verification. Install/use Playwright in the prototype workspace, launch Chromium headless, capture desktop and mobile screenshots, smoke-test the primary interaction, collect console errors, and then inspect screenshots with vision. This is a QA fallback pattern, not a claim that browser tools are broken.
   - For small-group decision products, use the reference `references/group-consensus-decision-prototypes.md`: build a decision room rather than a generic poll; make async participation, threshold/deadline rules, non-voter handling, hard-constraint filtering, least-regret scoring, and the shareable consensus explanation visible in the first artifact.
   - If a group consensus prototype already exists and Antoine rejects the design quality, use `references/dinnervote-v1-reset-pattern.md`: rebuild around the decision-room surface itself, not the old landing-page skeleton.
   - If Antoine asks for the rebuild specifically through Claude Code, use `references/dinnervote-claude-v2-rebuild.md`: complete CLI auth/use, handle partial Claude timeouts with focused continuation prompts, and independently verify the artifact before reporting.

7. If UI/UX feedback is sharply negative, run a salvage redesign pass
   - Treat “the UI/UX is horrible” as valid product evidence, not taste noise. Do not argue or explain the intent; rebuild the interface around the core customer promise.
   - Diagnose the failure in concrete terms before editing: clutter, weak CTA hierarchy, abstract hero copy, fake/prototype smell, dead empty state, parser-demo feel, missing long-term product object, insufficient trust for sensitive data, or product proof that looks static/toy-like.
   - If incremental CSS/button tweaks did not fix the reaction, stop patching the bad foundation. Take a screenshot, run brutally critical browser vision, list why the page fails commercially, then rebuild the composition around trust + proof + action.
   - When follow-up feedback prefers part of an earlier version, do not treat the latest redesign as sacred. Preserve the improved element (e.g. stronger H1) while selectively restoring the higher-signal earlier element (e.g. a concrete hero dashboard/image) in a cleaner system.
   - Re-anchor the page around a plain user promise, not an operating-system abstraction. Example for Homebase OS after failed versions: “Turn messy home bills and emails into one clean dashboard.”
   - For Antoine’s consumer prototypes, the hero should usually lead with product proof/interaction rather than a conventional landing-page pitch. This is especially true when the category is vague or trust-sensitive. The first viewport should answer “what exactly do I get and can I try it?” through the artifact itself: dashboard, digest, pod, source drawer, reviewed plan, or other concrete object.
   - HappyCouple positive reference: its hero works because a central AI pod creates a direct interaction target, “CLICK TO TALK” is attached to the object, the primary CTA reinforces the same action, the H1 is emotionally specific and not generic AI copy, the palette feels private/warm, and prompt examples reduce blank-page anxiety. Future prototypes need their own equivalent of that role, not a copied pod.
   - Use a credible consumer-product visual system: off-white/light background, black typography, restrained accent color, rounded product cards, readable tables/lists, and no neon/glowy AI styling unless the category calls for it.
   - Make the hero show the actual end-state product object, not just marketing copy: dashboard preview, monthly cost, next due item, changed item, open issue, saved docs, or equivalent persistent records.
   - If a section is meant to compare two concepts side-by-side or create a same-length rhythm, reduce content density and component height together: fewer records/actions, shorter text, smaller cards, tighter padding, and matching column heights. Do not just shrink fonts.
   - Do not create arbitrary asymmetric grids that make adjacent product/proof cards look randomly sized. Use a shared grid system (`repeat(2, minmax(0, 1fr))`, consistent gutters, shared container width) for side-by-side demo/input-output sections unless there is a deliberate visual reason not to.
   - Equal-height cards are only good when the internal content can justify the height. If one side is sparse (e.g. checkout/payment card), either make it a nested card inside one unified parent section, add meaningful content, or let it align to start. Do not stretch a sparse card into a tall empty rectangle just to match its neighbor.
   - Hash anchors and sticky navs can create fake layout defects: sticky headers may overlap section tops, making cards look clipped or oddly sized. For landing-page prototypes, either use `scroll-padding-top`/`scroll-margin-top` and test `#anchor` URLs directly, or remove sticky positioning if it harms the composition.
   - When the user provides a screenshot complaint about spacing/section size, analyze the screenshot first, identify the grid/anchor/content-balance failure, then fix the layout system. Do not keep patching typography or copy while the underlying grid is broken.
   - For sensitive categories (home docs, money, health, legal-ish, children/school/family logistics), put trust objections above the fold and near checkout: no broad inbox/account connection, no third-party contact, no acting/submitting/paying on the user's behalf, source-backed records, deletion/data-use framing, exact deliverable, refund/guarantee if valid, and what happens after payment.
   - For child/school/family-data products, do not ask for unlimited inbox access in the first test. Frame the safer wedge as selected forwarding/pasting, source-linked verification, explicit limits (“we never contact your school/teachers/portals/other parents”), optional initials/redaction, no model training on forwarded messages, and delete-on-request. The first business test is trust + willingness to forward real messages, not parser accuracy.
   - Make product proof source-backed. Show “Before” messy input and “After” structured output, with source snippets/confidence/actions so the dashboard does not look like a static mock or ChatGPT summary.
   - Give empty states structure. A blank output panel feels broken; use a muted dashboard skeleton plus a specific instruction.
   - Make prototype limitations honest internally, but do not expose customer-breaking notes like “local button simulates checkout” on public-facing screens. Track that caveat in final response/Obsidian instead.
   - Re-run tests/build/browser interaction and browser vision QA after the redesign. Only report back after the new UI is visibly no longer embarrassing and no serious visible blocker remains.

8. If asked to make prototypes “user-facing ready,” raise the bar from polished demo to invited-user readiness
   - Do not interpret “user-facing ready” as only visual polish. For any paid or sensitive-data demand probe, make the page acceptable for an invited stranger to see: clear offer scope, what happens after CTA/checkout, exact deliverable, delivery timing, support contact, refund/deletion policy, data boundaries, and operator/accountability copy.
   - If Antoine says “make it a real thing” or “we need to get real users,” stop polishing the landing page and build a production-probe loop: production server start command, deploy config, environment-driven Stripe Payment Link, persistent submission storage, post-checkout intake, dashboard/retrieval route, protected operator console, admin JSON/CSV export, `.env.example`, and a real-user launch SOP with invite copy and concierge fulfillment steps. This is not a full SaaS/account system; state the gap directly.
   - The thinnest real-user paid concierge probe should have: `npm start`/equivalent running the backend, `HOMEBASE_PUBLIC_ORIGIN`, `HOMEBASE_STRIPE_PAYMENT_LINK`, `HOMEBASE_ADMIN_TOKEN`, `HOMEBASE_STORE_DIR`, persisted checkout/submission records, `/admin.html`, `GET /api/admin/submissions`, and `GET /api/admin/submissions.csv` protected by `Authorization: Bearer $HOMEBASE_ADMIN_TOKEN`.
   - Admin/operator tooling must get browser QA too. An admin console that technically works but has poor contrast, hidden token inputs, unclear empty states, or default-link/nav styling is not acceptable for real-user fulfillment. Run harsh visual QA on empty state and, if possible, loaded rows; fix readability before reporting ready.
   - For sensitive submissions, operator exports should avoid dumping full private source text by default. Export IDs, timestamps, email, checkout id, home metadata, monthly cost/status, dashboard URL, and short previews; keep full source access behind the app/storage, not broad CSV sprawl.
   - Sensitive categories need trust copy both above the fold and near the conversion action. Browser-vision QA often passes layout but still blocks on trust/process/legal gaps; treat those as product blockers, not copy nits.
   - For child/school workflows, surface independent/not-school-affiliated positioning early, selected forwarding only, no inbox/school contact, no submitting/paying/sending on behalf of parent, sensitive-data warnings, and source-linked parent review.
   - For household/financial-doc workflows, surface no bank/Gmail/portal login, private upload after checkout, up-to-N source item scope, delivery timing, AI+human review disclosure if true, no model training if true, deletion-on-request vs self-service deletion accurately, refund window, and support email.
   - If legal/support links exist, verify they actually ship in production build and return 200 in the browser. Vite single-page builds may omit sibling `privacy.html`/`terms.html`/`refund.html`/`contact.html` unless `vite.config.js` has Rollup `input` entries for each static page. A footer link that works in source but 404s in `dist` is a launch blocker.
   - Do not claim checkout/upload/email is live unless it is wired. For invited/manual pilots, customer copy can say “we’ll send checkout/upload/forwarding instructions,” but final response and Obsidian must state that real Stripe/email/upload automation remains an external blocker before public traffic.
   - After final fixes, run a second browser/vision pass asking specifically for blockers before showing to an invited paid user, then verify legal pages with `fetch()` from the browser console.

9. Verify before reporting back
   - Run tests.
   - Run build.
   - Confirm local preview serves expected HTML.
   - Interact with the main CTA/sample path.
   - Fix obvious visual QA issues before final response.

9. Log the artifact in Obsidian
   - Save a prototype note under the relevant Consumer AI Factory opportunity/venture folder.
   - Update `00 Operating System/Factory Task Status.md` if the prototype changes project state.
   - Append the build log with action, why, artifact path, verification, decision, and factory learning.

## Implementation pattern

A good quick prototype can be:

- Vite/static frontend for speed.
- `src/<core>.js` for the first-value engine.
- `test/<core>.test.js` using `node --test` for fast TDD.
- `index.html`, `src/app.js`, `src/styles.css` for the demo UI.
- No backend unless the demand test requires real submission/storage.

For visual-generation feasibility spikes where output quality is the gating risk, use a local batch-evaluation repo instead of a frontend-first prototype. A reliable pattern is a small Python package with: `.env.example` for provider keys, a style/prompt catalog, provider abstraction, always-available mock provider, real hosted-provider clients when keys exist, timestamped `runs/` directories, `manifest.json`, `report.md`, contact-sheet/grid generation with Pillow, gitignored `data/input/` and `runs/`, and tests for prompts/styles/provider availability/manifest behavior. Verify with `python3 -m pip install -e '.[dev]'`, `pytest -q`, `python3 -m compileall -q src`, and an end-to-end mock batch before asking Antoine for API keys or real photos. When testing hosted image APIs, treat credentials and billing as part of the spike: store tokens only in ignored `.env`, never print or commit them, warn if a token was pasted into chat, run a tiny one-style smoke test before a full batch, inspect `manifest.json` before reporting, and distinguish provider-account blockers like Replicate `402 Insufficient credit`/`429 throttled until payment method` from code failures. If billing was just added, wait a few minutes and retry a single style first; if it still fails, check that credits/payment belong to the same account/workspace as the token before changing code.

If the user asks to “build it” after confirming the product does not exist behind the paywall, build the thinnest real fulfillment skeleton instead of more landing-page polish:

- Add tested checkout-start/domain primitives: create checkout id, persist metadata, return Stripe Payment Link redirect when configured, and return a local success redirect only for development.
- Add a post-checkout success/source-submission page: payment confirmed, visible order reference, email for delivery, source-material instructions, privacy/data-use reassurance, and one clear submit CTA.
- Add a persisted dashboard/submission artifact: save submitted source material and generated dashboard JSON locally or server-side so the flow produces a retrievable object, not just localStorage intent.
- Add a dashboard retrieval page/route keyed by dashboard id. It can be ugly-first, but it must prove the paid path creates something durable.
- Keep Stripe wiring environment-driven (e.g. `HOMEBASE_STRIPE_PAYMENT_LINK`). Never hardcode keys or fake that payment is live when the env var is missing.
- QA the post-payment page separately from the landing page. Trust blockers there are different: missing “payment confirmed,” no order reference, fake prefilled sensitive data, weak privacy reassurance, and ambiguous submit button labels.
- Do not prefill the paid handoff textarea with realistic personal/financial example data. On the landing demo it is useful; after checkout it looks like someone else’s private data and damages trust.
- If local long-running Node servers crash or silently exit on Antoine’s Mac, keep domain logic in testable JS but use a Python `http.server`/small Python route fallback for local serving, or verify backend routes with an in-process server test rather than pretending a flaky dev server passed.

Example verification commands:

```bash
cd /path/to/prototype
npm install
npm test
npm run build
```

Static preview fallback:

```bash
cd /path/to/prototype/dist
python3 -m http.server 3020 --bind 127.0.0.1
```

If Hermes background servers produce no logs or fail to bind, launch detached via Python and verify with `urllib.request` + `lsof`.

## Related references

- `references/body-neutral-progress-photo-assets.md` — visual/icon guidance for body-neutral progress-photo coach probes, including asset prompts and the rule to start with one or two QA'd assets before generating longer sequences.
- `references/player-agent-outreach-email-mvp.md` — trust ladder for player/career-agent outreach products: start with player-approved copy/mailto and reply-paste handling, then tracked alias, then Gmail/Outlook OAuth only after trust exists.
- `references/domain-name-and-registrar-checks.md` — naming and domain-availability workflow for demand probes, including GoDaddy automation blocks, Porkbun live-price checks, WHOIS caveats, and purchase-approval boundaries.

## Output standard

Final user response should include:

- prototype path,
- local preview URL if running,
- what changed,
- verification results,
- direct caveat if it is rule-based/not real AI yet,
- next demand-probe step.

## Pitfalls

- Do not equate a prototype with a company. The next step is evidence, not features.
- For agent products that act through a user's identity/email (player outreach, career agents, school/parent copilots), do not ask for full inbox access as the first activation step. Start with user-approved drafts, copy/mailto, mark-as-sent, and reply-paste handling; add tracked aliases and OAuth only after the user has seen enough value to trust the product.
- Do not make the page visually flashy when the category requires trust.
- Do not expose internal testing language to end users.
- Do not expose internal demand-measurement buttons like “I would pay for this”; use a believable user-facing conversion event instead, even if the backend is not wired yet.
- Do not confuse process improvement with artifact improvement. If Antoine asks whether “this UI” was updated, updating a playbook/skill/Obsidian note is not enough; change the actual app code, run tests/build/browser QA, and answer with the file paths and preview URL.
- If Antoine explicitly asks to use Claude Code / “do CLI” for a prototype rebuild, do not substitute a direct Hermes implementation and call it done. Authenticate/configure Claude Code if needed, run the rebuild through the CLI, then independently verify the resulting files with tests/build/browser QA. If Claude Code times out, inspect the partial artifact, run the failing verification yourself, and send Claude a focused continuation/fix prompt rather than restarting from scratch or reporting a half-run as complete.
- Do not confuse trust cues with a trust section. If Antoine asks to remove a “Trust layer,” remove the bulky standalone section, but lightweight reassurance chips near the hero/checkout may still be useful if they do not dominate the page. If he names exact copy to remove, remove that exact copy.
- Do not let the free preview cannibalize the paid test: make the free output a useful first pass, and make the paid offer clearly better through a genuinely valuable next step. Do not add human review unless the user’s trust problem obviously justifies waiting.
- Do not ship a demand-probe page without a conversion flow. A polished demo is not enough. The page needs a believable user-facing path from value proof to action: hero CTA -> checkout/start API -> payment or local preview fallback -> source/upload/setup handoff -> generated artifact/dashboard -> confirmation/review controls. Avoid fake internal CTAs. If the desired signal is payment, make the primary CTA a checkout action, not a waitlist/private-beta form. Do not add unnecessary pre-checkout fields; Stripe/checkout can collect email. For low-friction paid probes, prefer one obvious price and CTA language tied to the product object. If the product is meant to be persistent, do not use one-off artifact language like “Build my first home file” unless the flow actually ends at a one-time report; use workspace/account/setup framing such as “Create my Homebase — $19” and “Create my Homebase workspace,” then ensure the post-checkout handoff, source submission, and dashboard all reinforce that a durable workspace was created.
- For pure waitlist demand probes, do not let Lovable/AI builders produce a sparse generic section stack where the CTA is buried after “How it works” and “What this isn’t.” The waitlist action should be visible in or immediately adjacent to the first product artifact, then repeated only if the page is long. Compress around: product artifact proof → why this is different → sharp early-access promise → waitlist. Avoid excessive vertical padding, filler feature cards, and overly soft CTA copy like “quiet testers” unless the category explicitly needs that tone; the user should understand what early testers get, when/why joining matters, and what happens after submission.
- If Antoine objects that a paid CTA sounds like buying a report rather than signing up, treat that as a product-architecture bug, not just copy feedback. Update the artifact so the first paid step creates a persistent workspace/home base/account-like object; update tests to prevent one-off report language from returning; and verify the conversion path still works from landing CTA through post-checkout setup. Be direct in the final caveat if auth/persistence is still prototype-grade: copy can become more accurate, but the underlying product architecture also needs to support the promise.
- Do not add refund/legal boilerplate by default to early paid demand probes. If Antoine asks to remove refund language, remove it everywhere public: landing page, post-checkout/source page, legal nav/footer, build inputs, server routes, generated dist output, and any stale page files. Add regression tests that public source files/build config do not expose refund copy or links, then rebuild and verify `/refund.html` is absent/404 if the page was removed. Keep any refund/deletion/support policy intentional and accurate, not template residue.
- For local prototype QA, make the conversion flow testable even when Stripe/backend persistence is absent, but keep the fallback clearly development-only. A good pattern is: frontend CTA calls `/api/checkout/start`; production redirects to the backend/Stripe URL; localhost catches checkout failure and routes to a local setup page with a synthetic checkout id; setup submission writes a generated artifact to `sessionStorage`; dashboard reads that local artifact if API fetch fails. This lets browser QA test the full product flow without pretending payment is live.
- Do not keep building after the first credible interaction exists. If Antoine says the landing page is good/pretty good, stop polishing and move to a traffic/telemetry demand test. The next useful work is usually: verify waitlist capture, install analytics, connect a credible domain, send 100–200 targeted visitors, and judge waitlist conversion before building upload/auth/AI.
- For waitlist demand tests, make PostHog/analytics privacy-preserving from the start: track `landing_viewed`, CTA clicks, option/focus selected, section-seen events, and `waitlist_submitted`, but do not send raw email addresses, body/photo data, or private user text. Use properties like `has_email: true` and selected goal/focus.
- If the page is hosted in Lovable/no-code and the agent cannot edit code directly, create/verify the analytics project yourself when possible, then give a paste-ready Lovable prompt with the project key, event list, no-PII rules, and verification criteria. After Lovable applies it, verify live event arrival with a test signup and the provider events API.
- If user skepticism exposes weak venture logic, update the artifact/status and change course rather than defending sunk-cost prototype work.
- If user skepticism exposes weak venture logic, update the artifact/status and change course rather than defending sunk-cost prototype work. For Homeowner Triage, the better reframe became “Home Admin OS”: a persistent dashboard for home costs, documents, reminders, repairs, suppliers, and decisions. This is stronger because AI advantage comes from ingestion + memory + organization + reminders + workflow, not one-off summarization.
- Do not call a subtle same-skeleton pass a “rebuild.” If the user asks for a ground-up UI reset, create a visibly distinct version, preferably as a separate route/file such as `v2.html`, so old and new can be compared side-by-side. A true reset should change composition, product surface, density, visual rhythm, and art direction — not just copy, CTAs, trust chips, card styling, or minor hierarchy.
- When reporting UI/prototype changes, always include the exact local or hosted preview URL that serves the changed artifact. If the work is only visible through `file://`/inline QA artifacts or local ports, say that plainly. Do not claim something is “showable” without giving Antoine a way to see it immediately.
- For broad “operating system” ideas, do not make the first prototype a generic text analyzer. Show the persistent system of record directly: location/home-type setup, summary cards, concrete object tables, due dates, missing setup, alerts, next actions, source chips, and visible reminder/action states. The UI should answer “what will live here over time?” not merely “what did the AI generate from this paste?”
- For location-aware products, separate the universal data model from the local layer. Universal: profile, recurring costs, bills, documents, due dates, repairs/issues, suppliers/contacts, renewals, alerts, tasks. Local: tax/fee names, ownership/rental terms, document expectations, default reminders, vocabulary, and jurisdiction-specific setup checklist. Use a sharp local example like London only as a wedge/demo, not as the category boundary.
- If the user says location coverage is too thin, do not add two generic options and call it solved. Run a structured taxonomy pass across many representative markets, then update both the UI options and the underlying presets/tests. For home-admin products, research each market’s recurring costs, local tax/fee names, tenancy/ownership documents, building/community fees, utility conventions, repair actors, renewal/deadline patterns, and native vocabulary.
- For demos that require user input, never start with a giant blank textarea unless the user explicitly arrived with their own material. Pre-fill the realistic sample input and render a useful output on page load; let users edit from there. “Try sample” should be a convenience/reset, not the only way to see value.
- For ADHD/executive-dysfunction “unstuck” products, do not assume the right hero is a decorative output preview card. Antoine rejected a right-side Unstuck Card preview because it distracted from the actual experience. Prefer a calm centered first viewport: judgment-free benefit H1, short explanation, input box above compact researched stuck-task suggestion chips, and one CTA to break the task into tiny steps that fits above the fold. Avoid accusatory H1 copy like “What are you avoiding?” if it can read as judgment. Use readable calm typography over expressive/editorial display type; if browser QA says the H1 still feels heavy, soften font family/weight/letter spacing rather than defending it. If Antoine says the product feels like a gimmick, do not add more tiny-step polish. Reframe around the real wedge: an AI body double / companion session that creates presence, motion, accountability, and a completion receipt. A stronger first-value loop is rescue prompt → one task → editable tiny step → live check-ins/timer → receipt → paid-intent/waitlist CTA. Browser vision should explicitly check whether the first viewport reads as “someone will sit with me while I do this,” not just “a planner made a list.” The stronger product object may be a Focus Plan or Companion Session, not a single card: after input, open a focused one-task completion page with Tiny Steps, inline editing inside the step list, current subtask, “make smaller/gentler” controls, calm timer/check-ins, in-progress/complete step states, and automatic advance. If the companion/session still feels gimmicky because it lacks real AI, persistence, memory, or retention, do not keep polishing mascot/voice/timer copy. Test the durable-value wedge directly with a receipt-level demand probe: control waitlist CTA versus a “Stuck Journal” / pattern-memory variant that infers one honest pattern from the current session and asks permission to remember it next time. Keep copy honest: say “One pattern from today,” not “your remembered pattern,” and do not imply persistent memory exists before auth/history/backend do. Use stable localStorage assignment for lightweight A/B testing, add category-based insight helpers, and measure memory CTA intent against the generic waitlist baseline before building account infrastructure. If a brand motif emerges (e.g. hourglass for Start Small), carry it through both logo and timer; make minutes prominent and seconds/details visually smaller. Browser QA should verify not just beauty but state clarity: side controls readable, current step not mislabeled as pending, CTA says “Start this step” rather than ambiguous “Start task,” timer numerals are readable, inline-edit inputs render inside Tiny Steps, and the receipt-memory variant communicates future memory without claiming memory already exists. Make how-it-works centered typography rather than boxed SaaS cards; keep FAQ as minimal collapsible accordions. For lock-in/freeze-computer ideas, use soft focus mode and clear escape rather than hostile browser/computer trapping, especially because the task may happen on the computer.
- For broad location-aware products, standardize user-facing location selectors. Use stable codes internally (e.g. ISO-like `GB`, `US`, `FR`) and clean readable labels externally (“United Kingdom,” “United States,” “France”). Avoid mixed labels like “UK,” “Other / not listed yet,” or country aliases in the UI. Keep aliases in normalization logic, not in the selector.
- For sensitive-family workflows like school email copilots, avoid customer-facing phrases that sound like internal product strategy: “demand probe,” “business test,” “local intent,” “Stripe not configured,” “object,” or “do not send traffic.” Use customer-facing framing like “private pilot,” “selected messages,” “parent-ready digest,” “source-linked verification,” and “private forwarding instructions.” Keep the direct business caveat in Obsidian/final response, not on the page.
- For sensitive-family workflow demos, a giant editable textarea plus “refresh” can look like a toy AI prompt playground. Prefer a before/after transformation and structured product object first; make editing sample messages optional, rename it in plain user language (“sample school emails”), and show restraint/ambiguity flags plus source references so the output feels operational rather than magical.
- For dashboard/homepage previews, if the actual product object is central to the promise, consider making the interactive dashboard full-width below a compact composer rather than squeezing the input and output side-by-side. A cramped dashboard makes the product feel weak even if the logic works.
- When browser QA says the prototype still risks feeling like a ChatGPT wrapper, add proof-of-system elements before adding more copy: saved-home framing, last-updated/history cues, source audit trail, reminder states, persistent object lists, and explicit ingestion modes such as forward email/upload PDFs/paste messages.
