Deep research you can follow — claim by claim, source by source.

10 reports · 3,403 sources read · 248 receipts attached

Most recently, we reported · 26 Aug 2026

The LLM Edge in Finance Is a Division of Labor

THE QUESTION

Where does an LLM actually add value in a financial system—numerical forecasting, unstructured-information understanding, macro interpretation, portfolio decisions, or quantitative research—and where do traditional systems remain stronger?

WHY IT MATTERS

Financial systems already combine text, market data, macro releases, research code, portfolio construction, and execution. Those stages do not carry the same consequences: an inspectable feature can remain advisory, while a portfolio decision can move capital.

THE ANSWER

The evidence supports a division of labor. LLMs look most useful with semantic, heterogeneous inputs and inspectable, reversible outputs; statistical, machine-learning, and deterministic systems remain the stronger default for numerical prediction, calibration, optimization, constraints, and execution. Each module therefore needs a matched control and an endpoint suited to its claim, with point-in-time data, equal search budgets, repeated runs, and executable costs. No independently reproducible matched system in this evidence set shows persistent value across every role, so authority should contract as an output approaches a capital decision.

The evidence behind this report

We read 278 sources and cited 23. Every citation ships with a receipt — open one:

The receipt for [1]
Academicarxiv.org
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
Why we searched

2025 2026 point in time LLM finance benchmark module ablation text factor allocator end to end strong baseline costs multiple testing

What the source establishes

Existing financial QA benchmarks treat SEC filings as flattened databases, focusing on isolated fact lookup and failing to capture the cross-document, cross-temporal, and cross-entity integration required for professional financial analysis. Fin-RATE introduces three task types—DR-QA, EC-QA, and LT-QA—to separately evaluate fine-grained reasoning, cross-company comparison, and longitudinal tracking, thereby disentangling different sources of model error. The benchmark uses a dual-model generatio

Cited in 2 places in this report
Click any marker — the receipt shows what we read and why we searched for it.
Recently published
All 10 reports →
25 Aug Can This Checkpoint Still Learn? Why retention and future learnability need separate tests in repeatedly trained neural networks. It frames a three-arm choice: continue, reset training state, or retrain from scratch. 10 cited · 319 read 24 Aug The Memory Was Right. The Decision Was Wrong. Long-running agents need more than accurate retrieval: they need a disciplined way to turn past events into current, scoped, independently supported inputs to action. 15 cited · 272 read 23 Aug How SPADE turns self-generated worlds into a bounded training loop A model can improve without human-written QA pairs or stronger-teacher demonstrations. SPADE shows where the missing supervision moves—and which controls must govern the result. 10 cited · 192 read 22 Aug How a Reasoner Learns When to Stop A full-parameter stop/refine policy learns from both repair and damage; sampling, search, and external verification remain isolated experiments around that control. 24 cited · 280 read
Topics

Every report is tagged by the ground it covers; each tag is a standing thread.