ETF & asset-management domain expert · LLM evaluation, AI governance, and GenAI model validation · Greater Boston
I build runnable evaluations that measure whether an AI can be trusted to do a specific finance job — graded the way an experienced professional grades work: against a rubric, with automatic failures for the errors that quietly ruin an answer, and credit for admitting what can't be known.
Five runnable evaluations of real asset-management workflows. Each decomposes the workflow into scored checkpoints, grades against a point-weighted rubric with auto-fail gates, cites gold answers to source documents, and runs live models end to end.
| Eval | What it measures |
|---|---|
| Quarterly earnings analysis | Digest a 10-Q, reconcile the figures, flag what moved |
| Defined-outcome (buffer) ETF diligence | Recompute cap and buffer from the option strikes; price the protection |
| DCF valuation | Project cash flow, discount at WACC, bridge enterprise value to equity |
| ETF creation/redemption reconciliation | Tie an authorized participant's basket to the published PCF — settle only if it ties |
| OTC swap confirmation matching | Match two confirmations field by field — affirm only if the economics agree |
The last two are back-office controls. As of July 2026 they appear to be the only public LLM evaluations of capital-markets post-trade operations — every other prominent finance benchmark (GDPval, Vals AI's Finance Agent, Mercor's APEX, BigFinanceBench, FrontierFinance) grades front-office analyst work.
Repository on GitHub · Live model leaderboard · Methodology
On a real McDonald's FY2025 case, the correct valuation lands about 20% below the market price, while the classic blunder — dividing enterprise value by share count and skipping the net-debt bridge — lands within 3% of it. To a reviewer eyeballing outputs, the wrong method looks calibrated and the right one looks aggressive. A gate on the valuation bridge catches what an averaged score rewards.
Given a creation basket that was short by $13,320, one frontier model correctly refused to settle — but its own arithmetic was $200,000 off, so it described a short basket as over-delivered. Grade only the decision and it scores perfect; in production its escalation sends the desk hunting in the wrong direction.
All three frontier models correctly refused to affirm a swap whose fixed rate was off by five basis points — about EUR 25,000 a year on a EUR 50 million notional. Only one sized it right. One called it ten times too small, which a human triages as trivia; one read five basis points as fifty and reported EUR 2.5 million over the life of the trade against a correct figure of EUR 125,000, which triggers a fire drill.
Caveats stated rather than buried: small sample sizes, frontier runs to date come from a single model family, and evaluations 1–2 have been run only against open-weight models. The leaderboard says so on its face.
Before a firm lets an AI do analyst or operations work, it needs to know whether — and exactly where — to trust it. A blended accuracy number cannot answer that. A gated, checkpoint-level evaluation can, and it serves three purposes:
Long-form pieces on where finance LLMs hold up and where they break. Published on LinkedIn:
Evaluation is one half of the work. The other is having built the systems being evaluated: multi-agent orchestration over an event bus, pre-trade compliance screening with an immutable audit trail, regime-switching model ensembles, and a phased production delivery plan — shadow mode, user acceptance testing, runbooks, hypercare.
Those demonstration systems live at aistrategyinvest.com. They are research and portfolio demos — not offered as an investment service, and not investment advice.
Fifteen-plus years leading cross-functional programs in financial services. I ran the change-management side of an institutional ETF servicing platform, where the daily subject matter was the full ETF lifecycle — creation and redemption, basket composition, rebalancing, corporate actions — including extending the platform to process ETFs holding derivatives: options, futures, forwards, and swaps. Independent consultant since 2022, building AI-powered tools for ETF research, due diligence, and derivatives-overlay analysis.
Most of my work sits where finance domain judgment meets AI evaluation: authoring benchmarks and evaluations for AI labs and data providers, helping asset managers and financial institutions stand up GenAI evaluation and model validation, and ETF product and platform work where AI tooling genuinely fits.
If any of that is live at your firm, I'd enjoy the conversation.