Dmitry Krutous

ETF & asset-management domain expert · LLM evaluation, AI governance, and GenAI model validation · Greater Boston

I build runnable evaluations that measure whether an AI can be trusted to do a specific finance job — graded the way an experienced professional grades work: against a rubric, with automatic failures for the errors that quietly ruin an answer, and credit for admitting what can't be known.

finance-llm-evals — an open benchmark suite

Five runnable evaluations of real asset-management workflows. Each decomposes the workflow into scored checkpoints, grades against a point-weighted rubric with auto-fail gates, cites gold answers to source documents, and runs live models end to end.

EvalWhat it measures
Quarterly earnings analysisDigest a 10-Q, reconcile the figures, flag what moved
Defined-outcome (buffer) ETF diligenceRecompute cap and buffer from the option strikes; price the protection
DCF valuationProject cash flow, discount at WACC, bridge enterprise value to equity
ETF creation/redemption reconciliationTie an authorized participant's basket to the published PCF — settle only if it ties
OTC swap confirmation matchingMatch two confirmations field by field — affirm only if the economics agree

The last two are back-office controls. As of July 2026 they appear to be the only public LLM evaluations of capital-markets post-trade operations — every other prominent finance benchmark (GDPval, Vals AI's Finance Agent, Mercor's APEX, BigFinanceBench, FrontierFinance) grades front-office analyst work.

Repository on GitHub · Live model leaderboard · Methodology

Selected findings

The wrong method can look the safest

On a real McDonald's FY2025 case, the correct valuation lands about 20% below the market price, while the classic blunder — dividing enterprise value by share count and skipping the net-debt bridge — lands within 3% of it. To a reviewer eyeballing outputs, the wrong method looks calibrated and the right one looks aggressive. A gate on the valuation bridge catches what an averaged score rewards.

A model can make the right decision with the wrong number

Given a creation basket that was short by $13,320, one frontier model correctly refused to settle — but its own arithmetic was $200,000 off, so it described a short basket as over-delivered. Grade only the decision and it scores perfect; in production its escalation sends the desk hunting in the wrong direction.

The same break can be sized wrong in both directions

All three frontier models correctly refused to affirm a swap whose fixed rate was off by five basis points — about EUR 25,000 a year on a EUR 50 million notional. Only one sized it right. One called it ten times too small, which a human triages as trivia; one read five basis points as fifty and reported EUR 2.5 million over the life of the trade against a correct figure of EUR 125,000, which triggers a fire drill.

Caveats stated rather than buried: small sample sizes, frontier runs to date come from a single model family, and evaluations 1–2 have been run only against open-weight models. The leaderboard says so on its face.

What it's for

Before a firm lets an AI do analyst or operations work, it needs to know whether — and exactly where — to trust it. A blended accuracy number cannot answer that. A gated, checkpoint-level evaluation can, and it serves three purposes:

Writing

Long-form pieces on where finance LLMs hold up and where they break. Published on LinkedIn:

Building, as well as evaluating

Evaluation is one half of the work. The other is having built the systems being evaluated: multi-agent orchestration over an event bus, pre-trade compliance screening with an immutable audit trail, regime-switching model ensembles, and a phased production delivery plan — shadow mode, user acceptance testing, runbooks, hypercare.

Those demonstration systems live at aistrategyinvest.com. They are research and portfolio demos — not offered as an investment service, and not investment advice.

Background

Fifteen-plus years leading cross-functional programs in financial services. I ran the change-management side of an institutional ETF servicing platform, where the daily subject matter was the full ETF lifecycle — creation and redemption, basket composition, rebalancing, corporate actions — including extending the platform to process ETFs holding derivatives: options, futures, forwards, and swaps. Independent consultant since 2022, building AI-powered tools for ETF research, due diligence, and derivatives-overlay analysis.

MBA, Babson College · BA Economics, Northeastern University · PMP, Project Management Institute

Get in touch

Most of my work sits where finance domain judgment meets AI evaluation: authoring benchmarks and evaluations for AI labs and data providers, helping asset managers and financial institutions stand up GenAI evaluation and model validation, and ETF product and platform work where AI tooling genuinely fits.

If any of that is live at your firm, I'd enjoy the conversation.

dkrutous@gmail.com · LinkedIn · GitHub