Known limitations

The strongest claims are deliberately narrow.

The demonstration shows an inspectable workflow. It does not establish investment performance, analyst adoption, or commercial value.

Small bounded sample

The Experimental Research Idea Engine reviewed eight frozen cases. Two produced publishable hypotheses and six failed safely; the 60-case phase remains blocked and unrun.

Human review pending

Independent evidence review and automated grounding checks are complete, but formal blinded human ratings have not been collected.

No workflow pilot

No live institutional workflow pilot has measured analyst time, recall, trust, edit burden, adoption, or willingness to pay.

No performance claim

No Sharpe-ratio, information-coefficient, hit-rate, signal-decay, return-prediction, or causal improvement has been demonstrated.

No commercial validation

The project has no validated customer-adoption or commercial-success claim. Interviews informed the design but did not establish product-market fit.

Public-safe evidence limits

The public repository uses permitted, frozen evidence and cannot reproduce the breadth of licensed institutional datasets.

Static public packets

GitHub Pages serves frozen packets rather than live inference. This improves auditability but does not demonstrate a production generation service.

Generalization risk

Results from the bounded issuers, filings, periods, and disclosure categories may not generalize to other companies, markets, document types, or regimes.

Analyst accountability

A hypothesis can be grounded and still be unhelpful or wrong. Analysts remain responsible for interpretation, additional diligence, and every decision.

Release boundary: formal blinded human review is the next evidence gate. Broader generation is not authorized.

Review the sequenced future research plan