# clone.arnao.ai — Product Requirements

*A recorder that learns you and becomes you — built by the person whose job is to make that safe.*

Owner: Byron Arnao (AWS Principal Technologist, Global Responsible-AI Lead)
Status: v1 PRD · noindex · front-door / brand-level scrutiny

---

## 1. Product & the wedge

clone.arnao.ai ingests you — live camera and mic, plus speech, video, photos, files, markdown, URLs, and running agents — and assembles a **persona model**: face, voice, cadence, knowledge, mannerisms, values. That model then shows up two ways. As a **realized clone** in a web interface that looks, sounds, and talks like the source. And as an **autonomous agent** that acts in your voice and values under a governance layer. Google auth, a guided clone builder, and one control that matters most: a **quality-vs-latency slider where realism is king**.

The market is already flooded with talking-head and voice-clone apps. Almost all of them treat consent and disclosure as a checkbox and a footer. That is the opening. **A clone product from the Responsible-AI lead wins by making responsibility the product, not the fine print.** Consent is captured on camera. Disclosure rides on every frame. Every claim the clone makes traces back to the source material that taught it. There is a kill switch the owner can hit from a phone. This is not compliance theater bolted to a deepfake engine — it is the feature set, and it is defensible precisely because the cynical incumbents cannot copy it without abandoning the shortcuts that make them cheap.

The name promises a *recorder that learns you*. The clone is not minted once and frozen. Each session adds signal; the persona model accretes and sharpens the way a relationship does. That arc — "it knows me better than last week" — is the emotional hook and the retention loop, and it rides directly on infrastructure Byron already runs.

The worked example sets the ceiling for taste: a video of Byron's grandfather becomes a clone that carries his warmth and his stories. Done with an estate gate and a memorial frame, that is legacy work families would pay for and trust. Done carelessly, it is the thing this product exists to prove unnecessary. We build the dignified version and make the guardrails the reason it feels safe.

## 2. Architecture

**Ingestion, per modality.** Each input runs its own extractor and reports into six fidelity meters — **face · voice · cadence · knowledge · mannerisms · values** — so the user watches the model fill in.

- *Camera (live/video)* → face geometry, expression range, micro-gesture and blink timing → face, mannerisms.
- *Mic / speech* → speaker embedding for Chatterbox, prosody, pacing, filler words, laugh → voice, cadence.
- *Files / markdown / URLs* → chunked, embedded, and stored as a cited knowledge base → knowledge.
- *Agents / endpoints* → behavioral samples of how the source decides and phrases → mannerisms, values.
- *Photos* → identity anchors and aging references (central to the grandfather case) → face.

Every chunk is stored with a **source pointer** (which file, which timestamp, which utterance). Provenance is not reconstructed later; it is captured at ingest and never separated from the derived trait.

**Persona model.** The assembled representation is deliberately **inspectable and editable**, not a black box — this is the interpretability stance the brand demands. Face and voice are embeddings. Knowledge is a retrieval index with citations. Cadence, mannerisms, and values are structured, human-readable trait cards ("understates achievements," "defaults to teaching") each linked to the evidence that produced it. The owner can read the model of themselves, correct a wrong trait, and delete a source — which recomputes anything that source touched. **Learning over time** is the same pipeline run incrementally: new sessions append signal, and the meters climb, backed by the fleet's existing session-to-session recall.

**Rendering — the web surface.** Chat drives a video-avatar face lip-synced to **Chatterbox** voice output, streamed to the browser. Retrieval grounds every answer in the knowledge index so the clone speaks from the source, not the base model's guesses.

**Running as an agent.** The same persona model is loaded by an agent on Byron's fleet. Its knowledge, voice, and value cards become the system context; its permissions and kill switch come from the governance layer in §3. One model, two surfaces.

**Where the slider trades.** Realism is king, so the default favors quality; the slider exposes the cost honestly.

- **Cinematic** — full cloud render, highest-fidelity avatar and voice, seconds of latency. For legacy pieces and set-piece conversations.
- **Balanced** — cloud voice, lighter avatar, pre-rendered openers and cached frequent answers to hide the round trip.
- **Instant** — on-device/lightweight avatar, fast voice, retrieval trimmed. For quick back-and-forth.

The real levers are cloud-vs-local render, a **cache** of common exchanges, and **pre-rendering** predictable lines (greetings, signature stories) so perceived latency drops without lowering the ceiling on quality.

## 3. The responsible-clone layer (first-class)

Every mechanism below is visible in the UI. A clone that cannot show consent, disclosure, and provenance is a defect, not a draft.

- **Consent capture.** A living source records consent *on camera* — an utterance naming the clone, its scope, and its duration — stored with the persona model as a verifiable artifact. No consent recording, no build.
- **Deceased / estate provenance (the grandfather).** A distinct, heavier gate: a documented family/estate authorization plus a memorial framing surface that states, plainly, this is a remembrance built by his family. The relationship of the requester to the deceased is recorded. The frame is memorial, never impersonation.
- **Disclosure by default.** A persistent **"AI clone" ribbon** on the web surface and a spoken/first-line disclosure on the agent surface. It cannot be dismissed. The clone never passes silently as the real person.
- **Provenance per claim.** Hover any statement to see *how the clone knows this* — the exact source (video timestamp, file, utterance). Unsourced assertions are visibly flagged as inference.
- **Kill switch + scoped permissions.** The owner freezes any clone instantly from the web or phone. Agent permissions are explicit and legible: *may speak as me* / *may draft* / *may NOT transact, spend, or make commitments* by default.
- **Watermark + audit trail.** Rendered media carries **C2PA content credentials** (provenance metadata signed into the file); every clone action — web reply or agent step — is logged with timestamp, surface, and the permissions in force.

## 4. Reuse of Byron's existing infra

This ships fast because most of the hard parts already run on the fleet.

- **Voice = local Chatterbox** on the Mac host (not ElevenLabs), reusing Gia's **`clone` skill** and the existing **`mtl-clone.py`** multilingual path — which already enforces a **mandatory per-language disclosure line** and a pronunciation-lexicon / whisper-back QA pass. The responsible-voice pipeline is already Byron's standard; the clone inherits it.
- **Auth** follows the proven **Google auth pattern** from the other arnao.ai apps — no new identity stack.
- **The agent surface** runs on the **existing agent fleet** with its scoped-permission and kill-switch primitives, tying directly into the harness/agent-governance thesis Byron already argues.
- **Storage** uses **Vercel Blob** for media and persona artifacts; **deploy via the vercel-deploy skill** to `clone.arnao.ai`, noindex, per front-door rules.
- **Memory / accretion** reuses the fleet's session-to-session recall so "learns you over time" is configuration, not new research.

## 5. Honest risks (the brand-defining section)

The point of this product is that the person building it names the risks first and out loud.

- **Likeness & legal.** A clone is a person's identity. *Mitigation:* on-camera consent for the living, an estate gate for the deceased, a right to delete that recomputes the model, expiry/renewal on consent, and C2PA provenance so any output can be traced to its authorized origin.
- **Grief-tech ethics.** A remembrance clone can help mourning or exploit it. *Mitigation:* memorial framing, not impersonation; family-only provenance; refusal to fabricate — the clone speaks from sourced material and marks inference as inference; no upsell against grief. The grandfather case is the reference standard for taste, and it is opt-in, family-gated, and disclosed.
- **Misuse / fraud.** Clones are a scammer's dream. *Mitigation:* undismissable disclosure, agents that **cannot transact by default**, watermarking, a full audit trail, and an instant kill switch. The default posture is "cannot do harm," and every loosening is an explicit, logged owner choice.
- **Over-trust of the model.** A confident clone can be confidently wrong. *Mitigation:* the persona model is inspectable and correctable; provenance-per-claim keeps the human able to check the clone against its sources.

These are not disclaimers. They are the argument for why this clone, from this author, is the one to trust.

## 6. Roadmap & recommendation

**Phase 1 — The Clone Builder.** Ingestion UX, live fidelity meters, consent capture, provenance tagging, Chatterbox voice, the inspectable persona model. Ship the wedge.
**Phase 2 — Face to Face.** The realized web clone: video-avatar, chat, disclosure ribbon, provenance-on-hover, the quality/latency slider — and the dignified grandfather build.
**Phase 3 — Clone-as-Agent.** The governed autonomous surface: scoped permissions, audit log, kill switch on the fleet.
**Phase 4 — Persona interpretability.** Deep inspect-and-correct tooling and the accretion timeline made visible.

**Build Direction 1, the Builder, first — for two reasons.** It is the **dependency root**: no Builder means no clone to face, no agent to govern, no model to inspect. And it is where the **moat is instantiated** — consent, provenance, disclosure, and the fidelity meters all originate at ingestion, so the Builder is the single direction that most fully *embodies* the differentiator. Face to Face is the more emotional demo, but it presupposes a built clone; it is Phase 2, not first. Start where trust is manufactured.
