Digital Scholar
Back
Digital Twin Methodology

How our expert digital twins are built

Every conversational digital twin on this platform is built from the same stack of published techniques — not model tricks. This page names each technique, describes how we apply it, and cites the paper, standard, or documented source of the approach so you can verify the underlying method independently.

Read the 3,000-word white paper
01

Conversational Digital Twin framing

Best practice

Model the expert as a persona layer over a corpus, not as a fine-tune. Keep persona and knowledge separate so both can be edited and audited.

How we use it

Each expert has a `system_prompt` (persona / voice / rules) plus a private, chunked, embedded corpus. Grounding rules are appended by the app so the persona editor can't accidentally disable them.

Precedent & sources
02

Retrieval-Augmented Generation (RAG)

Best practice

Instead of relying on a model's parametric memory, fetch relevant passages from a trusted corpus at query time and instruct the model to answer only from them, with inline citations.

How we use it

User query → embed with `google/gemini-embedding-001` (3072-d) → cosine-nearest 6 chunks from pgvector → passages injected into the system prompt with numeric IDs → model is told to answer only from those passages, keep source markers internal, and refuse when uncovered.

Precedent & sources
03

Semantic vector search with pgvector

Best practice

Store document embeddings in a purpose-built vector index; use approximate nearest-neighbour (HNSW / IVFFlat) with cosine distance for sub-second retrieval at scale.

How we use it

Chunks are stored in Postgres `vector(3072)` with an HNSW cosine index. Retrieval is a stable-security-definer SQL function (`match_chunks`) scoped per-expert so corpora are isolated.

Precedent & sources
04

Confidence banding from retrieval similarity

Best practice

Surface uncertainty to end users. When the top-retrieved passage is a weak match, the answer is more likely to be extrapolation, not grounded expertise — signal that visibly.

How we use it

The best cosine similarity from `match_chunks` is bucketed into High (≥ 0.72), Medium (≥ 0.58), Low (< 0.58) and rendered as a pill next to every answer. Bands are calibrated for gemini-embedding-001.

Precedent & sources
05

Shadowbox / self-critique second pass

Best practice

A second model call, acting as a skeptical critic, checks the draft against the source passages, lists issues, and revises. Reduces hallucination and unsupported claims.

How we use it

After the primary generation, `critiqueAndRevise` prompts gemini-2.5-flash to return strict JSON `{issues[], verdict, revised}`. If `verdict === "revise"` we send the revised answer, and store both draft and critique in the audit log.

Precedent & sources
06

Independent-judgment audit log

Best practice

For expert-witness and regulated use, every AI-generated statement needs a defensible provenance trail: what was asked, what sources were used, what the model drafted, what a human reviewed and adopted.

How we use it

Every reply writes an `answer_audits` row with question, retrieved chunks, draft, critique, final reply, confidence, model, and timestamp. Admins review each answer in `/admin/review` and mark it Adopted, Reviewed, or Rejected — creating the record a Daubert challenge expects.

Precedent & sources
07

Fidelity Score (1–10 human rating)

Best practice

Turn the human review queue into a real feedback signal by asking the reviewer to rate each answer on a numeric fidelity scale, not just Adopted/Rejected. The score becomes a trend line for corpus quality, persona drift, and model regressions.

How we use it

Every answer in `/admin/review` gets an optional 1–10 fidelity score alongside Adopt / Reviewed / Reject. The dashboard shows the running average across all rated answers so a drop in fidelity is visible immediately — the same metric used in LLM eval harnesses like G-Eval and HELM.

Precedent & sources
08

Tamper-proof audit chain (hash-linked ledger)

Best practice

For expert-witness, regulated, and legal use, the audit log itself must be defensible: any edit or deletion of a past answer must be detectable. Best practice is an append-only, hash-chained ledger — each row's sha256 hash depends on the previous row's hash, so altering any row invalidates every row after it.

How we use it

A BEFORE INSERT trigger on `answer_audits` computes `row_hash = sha256(prev_hash ‖ canonical(row))` and refuses updates to any hashed column or deletions of any row. Reviewers can freely add review status, notes, and fidelity scores — but the question, sources, draft, critique, and final reply are frozen forever. `verify_answer_audits_chain()` replays the whole chain on demand and returns the first broken row, which the admin console exposes as a one-click 'Verify chain' button.

Precedent & sources
09

Persona + refusal rules

Best practice

Separate 'who the expert is' from 'what the expert knows'. Persona lives in the system prompt; refusal-when-uncovered is enforced by app-owned rules so it can't be edited away.

How we use it

`buildExpertSystemPrompt` composes persona + mode-specific rules (text vs. voice) + the retrieved SOURCE PASSAGES. The rules mandate first-person voice, no visible source-file names or bracketed citations to users, answer-only-from-passages, and a polite refusal + related-topic suggestion when passages don't cover the question. Source markers remain internal for admin audit.

Precedent & sources
10

Faithful voice cloning + fast STT

Best practice

For a digital twin to feel like the expert, use the expert's own cloned voice, low-latency streaming TTS, and low-error transcription. Preserve consent and revocation controls.

How we use it

Speech-to-text: `openai/gpt-4o-mini-transcribe` via the Lovable AI Gateway. Text-to-speech: ElevenLabs `eleven_flash_v2_5` with a per-expert voice ID (falling back to a default). Voice IDs are stored per expert so a twin can be re-voiced or revoked.

Precedent & sources
11

Client-side voice activity detection

Best practice

Push-to-talk is friction. Energy-based VAD with pre-roll and a silence hangover gives natural, hands-free turn-taking without shipping a full-duplex realtime stack.

How we use it

The browser captures 16 kHz mono PCM, runs an adaptive RMS-based VAD with a 1500 ms silence hangover, an adaptive noise floor, and a 6-chunk pre-roll. The audio pipeline trims trailing silence, downsamples to 16 kHz, and transcribes with OpenAI GPT-4o-mini-transcribe for low-latency STT before the server responds.

Precedent & sources
12

Privileged workspaces (Case Rooms)

Best practice

For real work — legal matters, board disputes, M&A, patient files — end users must be able to bring private documents to an expert without exposing them to other users or leaking into the public corpus. Each engagement is its own siloed workspace with its own retrieval namespace and its own audit trail.

How we use it

A signed-in user creates a Case Room bound to one expert. Files upload to a private `matter-docs` storage bucket keyed by `matter_id`. Embeddings go into `matter_chunks` scoped by `matter_id` with a dedicated `match_matter_chunks` SQL function. At query time we retrieve from BOTH the private matter and the expert's public corpus, tag every citation as CASE FILE vs. EXPERT CORPUS, and prioritize private context. RLS restricts view/write to the owner and site admins only; nobody else — including other paying users — can see the files, messages, or embeddings.

Precedent & sources
13

Row-level security & per-expert corpus isolation

Best practice

Multi-tenant expert data must not leak between experts or users. Enforce this at the database, not in application code — RLS + security-definer retrieval functions.

How we use it

All app tables have RLS enabled. `match_chunks` and `get_expert_for_chat` are `SECURITY DEFINER` with `SET search_path = public`, scoped by `_expert_id`. Roles live in a separate `user_roles` table checked via `has_role()` — never on the profile.

Precedent & sources
14

Folder → Synthesis → Twin (canonical knowledge pipeline)

Best practice

Raw source material — transcripts, articles, decks, emails, book chapters — is noisy, redundant, and versioned inconsistently. Ingesting it directly produces bloated retrieval, weak citations, and unauditable knowledge. Best practice is a curated middle layer: raw sources live in topical folders owned by the expert's team; an AI synthesis step compresses each folder into one canonical markdown summary; only the summaries are embedded into the twin. Updates re-synthesize a single folder, not the whole corpus.

How we use it

For Russ Rosenzweig, knowledge files follow the grouped naming convention: `persona_00_…` through `persona_02_…` for the persona layer (voice, FAQ, refusals, case vignettes) and `domain_00_…` through `domain_29_…` for the domain layer (admissibility, Daubert, testimony, ethics, etc.), totaling 33 files. Each file begins with YAML front-matter (title, category, expert, version) parsed into a `documents.metadata` JSONB column, and the body is chunked with heading-aware boundaries. Retrieval is balanced across categories so no single file dominates. When Russ's staff updates a file, only that file is re-chunked and re-embedded. This separates raw source (Drive) from curated knowledge (twin) from persona (system prompt) — three independently editable layers. This is Knowledge Base Methodology v1.2.

Precedent & sources

This methodology is deliberately conservative: every AI answer is grounded in the expert's own corpus, checked by an adversarial second pass, banded by confidence, and logged for human review. Where we depart from best practice we say so on this page — and if you find a citation that's out of date, tell us and we'll update it.