IRC §1091 Wash-Sale Calculation Drift
Failure Mode: During multi-lot tax-loss harvesting evaluation, generative model outputs hallucinated 30-day lookback replacement lot dates, creating false tax savings.
Designing AI trust in high-stakes finance.
An AI-assisted portfolio terminal where the hardest problem isn't the algorithm — it's calibrating human trust. How do you show non-deterministic AI predictions next to hard margin, leverage and tax math without inviting misuse or manufacturing false certainty?
Executive Architecture Decision Record
Advisors and prop traders reject conversational AI "black boxes" in regulated portfolios due to unquantified model risk (SR 11-7) and catastrophic liability on hallucinated tax/margin calculations.
Option A: Autonomous AI Advisor (Rejected: fiduciary breach).
Option B: Chatbot with Disclaimer Footnote (Rejected: 18% adverse error rate).
Option C: 5-Dimension Confidence Decomposition + Heat Meter (Chosen).
Adopted Option C: Zero AI-computed margin math. The LLM only narrates server-audited deterministic calculations; confidence spreads >30 trigger mandatory human escalation.
Achieved 95% unaided threshold recall in n=40 study; reduced adverse volatility errors from 18% to 4%, providing a fully compliant audit trail for compliance officers.
The contribution · at a glance
An independent concept that tackles the hardest UX problem in fintech: presenting non-deterministic AI predictions alongside hard margin, leverage and tax math without inviting misuse or creating false certainty. Figures describe design scope and a recruited-participant study — not shipped business outcomes.
Modelled-vs-measured: n=40 mixed-methods (20 RIA + 20 prop-desk · 120 sessions · latin square · think-aloud + NASA-TLX), 11.4s→4.1s median, and 38/40 unaided threshold recall are study-measured. 5 surfaces, 6 personas, 12 citations, and 0 AI-computed margin are design scope. Independent concept — not a shipped product. Not an A/B test.
The challenge
Retail investors lack institutional-grade decision support, making portfolio management emotional and reactive. But integrating AI into money decisions creates three failures at once — and the interface has to defuse all three.
Presented as “definitive advice,” AI tested with 100% abandonment after the first wrong prediction. The trust curve was binary — total confidence, then total rejection, no graceful degradation.
Margin, leverage and tax live on separate screens, forcing users to synthesise by hand. Under market stress that produces emotional, high-pressure errors — exactly when precision matters most.
In ASIC/SEC-regulated markets, AI output that reads as “financial advice” creates legal exposure. Fact vs. prediction has to be unmistakable — through visual language, not a buried disclaimer.
Decision framework · handling AI uncertainty
I evaluated three approaches against user-trust metrics. Trust is calibrated through transparency — not suppression, and never through false confidence.
AI as definitive advice. Extreme early engagement, then 100% abandonment on the first inaccurate call. Binary trust: total confidence → total rejection.
AI kept fully separate from the calculators. Legally safe, but users cross-referenced across tabs — adding +2 min task time and defeating the purpose of integration.
AI inline with calculator output, labelled probabilistic: “if X, your margin exposure might be Y (78% confidence).” Visible confidence intervals, drill-into-reasoning. Trust calibrated through transparency.
Live · Co-Pilot
A single AI verdict is the wrong primitive for institutional decisions. Nova routes every query through six investor personas, each emitting a score with a 95% confidence interval. Divergence is the signal: when Buffett says 72% and Soros says 31%, both numbers stay visible — the PM reads the spread, not the mean. Click any persona to expand its reasoning trace.
Scores update live · 95% CI shown as the band around each dot · when spread > 30 pts the query auto-escalates to a human analyst before the PM sees a final answer (SR 11-7 model-risk routing).
Live · trust architecture
The failure mode of every AI output in finance is the confident-sounding answer that's wrong. Nova decomposes confidence into five independent dimensions — so a PM can see why the model is or isn't certain. A high overall score with low temporal relevance is a very different risk than a uniformly moderate one. Drag the handoff threshold; click any dimension.
Interactive · Heat Meter
Margin is a regulatory problem before it's a UI problem. The Heat Meter bakes the three real thresholds directly into the gauge — Reg T 50% initial, house 30% cushion, FINRA 4210(c) 25% maintenance. A position at 28% equity isn't “yellow” — it's 3 points above the maintenance call, 2 below the house warning. Run the stress test.
Process & evidence · 5-week study
Counter-balanced latin square, n=40, six risk-read tasks per session, think-aloud + NASA-TLX. Each variant won at least one metric — the question was which trade-off a regulated product could afford.
| Metric | V1 · Table | V2 · Gauge | V3 · Heat Meter |
|---|---|---|---|
| Time to decision (median, s) | 11.4 | 3.2 | 4.1 |
| Threshold recall (Reg T + FINRA) | 22 / 40 | 0 / 40 | 38 / 40 |
| Error rate (adverse-vol task) | 18% | 22% | 4% |
| NASA-TLX load (0–20, lower better) | 13.8 | 6.1 | 7.4 |
| Explain-to-client confidence | 3.4 / 5 | 2.8 / 5 | 4.6 / 5 |
The gauge is faster. V2 wins raw speed — but 0/40 participants could recall the Reg T or FINRA 4210 threshold afterward. V3 gives up 0.9 seconds and buys back 95% threshold recall and adverse-vol accuracy — the decisive metrics for a regulated product. A compliance-aware product doesn't optimise glance time in isolation.
What we cut
Each was reasonable in isolation, each failed against an institutional use case we later reproduced in testing. Naming them is part of the audit trail, not the marketing copy.
V2's giant centre percentage dominated everything — traders anchored on it and stopped reading. V3 puts the number inline with the threshold string (“38% · 8 pts above FINRA call”) so it can't stand alone.
Colour-only encoding broke for the 4.5% of participants with red-green CVD and for the Japanese cohort's amber salience. V3 uses four named bands with text labels + ARIA announcements.
An early build let the LLM compute margin for exotic cross-pairs — it produced plausible but impossible numbers (negative maintenance margin). V3 computes margin deterministically; the LLM narrates why, never what.
Product surfaces
Every surface answers the same question with different data: is this action safe to take right now? These are the real terminal screens — live Yahoo Finance quotes, FRED macro chips, a 12-holding institutional seed portfolio, Kelly + wash-sale + margin math all client-side.




Regulatory provenance
Nothing on Nova's surfaces is a design guess. Every number, default and flag traces to a specific regulation, rulebook or paper — hiring managers can audit the reasoning; analysts can defend the outputs under review.
Three families do the work: margin rules (Reg T 50 · house 30 · FINRA 25) draw the Heat Meter's named corridors; model-risk guidance (SR 11-7) decides when persona divergence must escalate to a human; disclosure and tax rules (ICA §5(b)(1) · IRC §1091) decide what the surface must say before a user acts.
| Tab | Surface element | Threshold / rule | Citation |
|---|---|---|---|
| 01 Co-Pilot | 95% confidence interval | Institutional disclosure band | NIST/SEMATECH §1.3.5.2 |
| 01 Co-Pilot | Persona divergence escalation | Spread > 30 pts → human review | SR 11-7 model risk |
| 02 Portfolio | Concentration KPI (top-3) | Diversification disclosure | Investment Company Act §5(b)(1) |
| 03 Heat Meter | Initial margin 50% | Regulation T initial requirement | 12 CFR §220.12 |
| 03 Heat Meter | Maintenance margin 25% | FINRA minimum maintenance | FINRA Rule 4210(c) |
| 03 Heat Meter | House margin 30% | Broker-discretionary cushion | FINRA Rule 4210(e)(8) |
| 04 Risk | Half-Kelly default | Institutional risk-reduction factor | Thorp (1997) · Poundstone (2005) |
| 04 Risk | Five-question risk quiz | Suitability assessment | FINRA Rule 2111 |
| 05 Tax | Wash-sale 30-day window | Disallowance on substantially identical | IRC §1091(a) · Pub. 550 |
| 05 Tax | Short/long-term boundary | Holding period > 1 year | IRC §1222(3) · Pub. 550 |
| 05 Tax | Specific-ID lot selection | Identification of sold securities | Treas. Reg. §1.1012-1(c) |
Engineering notes
A designer who can't hand off to engineers is a sketcher. Here's how every surface on the live terminal is wired — auditable end-to-end, not just the visuals.
Zero build-time deps. Three files — index.html 39KB, style.css 43KB, script.js 44KB. BEM .nv-* namespace, token-first (--nv-accent:#c9a959). ~126KB first paint.
Quotes from query1.finance.yahoo.com via corsproxy.io, 5-min refresh under rate limit. Graceful fallback to simulated data — the terminal never stalls on “Loading…”
TradingView's open-source renderer — the same library behind Binance and Interactive Brokers web. Crosshair, tooltip and time-scale interactions are ship-grade; no re-invention.
Every financial computation runs in-browser — no server round-trip, no data leaves the session. Kelly, FIFO/LIFO lots, §1091 window detection, Reg T / FINRA corridor math — all auditable in script.js.
Tabs as role="tablist" with ←/→ nav and aria-selected sync. All controls labelled. prefers-reduced-motion disables animation. Deuteranopia-safe status palette.
Every image carries width + height so layout shift is zero. State is in-memory — session ends, state ends. No cookies, no trackers, no fingerprinting payload.
Model Risk Post-Mortem
High-stakes AI must survive adversarial market edge-cases. Two incident post-mortems from stress evaluation.
Failure Mode: During multi-lot tax-loss harvesting evaluation, generative model outputs hallucinated 30-day lookback replacement lot dates, creating false tax savings.
Failure Mode: High synthetic sentiment scores masked severe bid/ask spread risk, misleading retail beta-testers into taking outsized leverage.
Multi-dimensional impact
All figures from the five-week study, n=40. Concept validation, not production business outcomes — each number carries its source.
Trust here was measured as behaviour, not sentiment: threshold recall asks whether users know where the regulatory line is; explain-to-client confidence asks whether they can defend the position; adverse-volatility error rate asks whether the surface prevents miscalibrated action. Speed counts only because those three held.
Users don't need the AI to be perfect. They need to know exactly when it might be wrong.
Ed Chen · Senior Product Designer · transparency is the highest-converting feature in high-stakes AI
Explore further
Every problem we solve for clients has multiple valid approaches — different costs, different ROI, different risk profiles. These threads show how the approach on this page compares to others in the portfolio.
Portfolio-level math primitives — HHI, beta, VaR, regime — rendered into UI defaults and AI-assisted decision surfaces.
Luxury, editorial, and brand discipline applied to financial interfaces — where restraint itself is signal.
How upstream regulation and macro prints become downstream product defaults and Legal-safe disclosure.