Empirical journey

The market hypothesis narrowed. The workflow opportunity became clearer.

Three locked studies separated what the evidence supports from what remains unproven.

Negative and null results were preserved rather than tuned away.

Research in plain language

Capability is supported. Market and commercial claims are not.

This is the shortest accurate interpretation.

Supported
  • Reliable document processing
  • Only information available before the filing was used
  • Evidence-linked change detection
  • Historical comparison
  • Reproducible classifications
Not supported
  • Incremental news alpha
  • Improved expected-language reconstruction
  • Causal claims
  • Commercial adoption or willingness to pay
Inconclusive
  • Validated downside-risk prediction
Not tested
  • Analyst time savings, trust, recall, and repeat use

Research decision tree

Modeling stopped when each evidence gate failed.

The null results did not end the project; they prevented unsupported claims.

Trading hypothesisDoes added document data improve prediction?
Gate 1 · not supportedStop incremental-alpha claim
Expected-language modelDoes extra data beat prior 10-K?
Gate 2 · not supportedStop added model complexity
Workflow productTest analyst value next

Study 1 · combined multisource model

Did 10-K information improve the 8-K and inexpensive-news model?

Question
Did adding 10-K features add one-month return information?
What was tested
Locked historical 8-K, 10-K, news, company-characteristic, and return panels with time-ordered folds.
Result
Adding the 10-K information did not improve the model once ordinary company characteristics were accounted for.
Decision
Not supported Do not advance an incremental-alpha claim.
Lesson
Attractive absolute statistics were not evidence that document information added value.
Technical metrics

Combined mean monthly Spearman IC: 0.0862. Incremental IC versus 8-K plus EODHD: −0.00227; one-sided p = 0.9531; two of five folds positive; 95% block interval [−0.00567, 0.00049]. Decision label: NOT SUPPORTED—EXPLORATORY HISTORICAL LOCKED.

Not claimed: alpha, causality, executability, live trading performance, or incremental combined-news information.

Study 2 · expected-disclosure model

Could other data predict the next 10-K better than the prior 10-K alone?

Question
Did company characteristics, intervening 8-Ks, and news improve the next filing’s representation?
What was tested
Consecutive filing representations and strictly prior characteristics, 8-Ks, and headlines.
Result
The prior 10-K alone predicted the next filing representation better than every augmented model.
Decision
Not supported Do not advance an enriched expected-language shock.
Lesson
Direct historical comparison was more defensible than added modeling complexity.
Method detail

All features obeyed point-in-time eligibility. The persistence benchmark won the locked comparison.

Not claimed: that news or characteristics add predictive information beyond the prior filing, or that representation error predicts returns.

Study 3 · disclosure-change demonstration

Can the system identify and evidence meaningful filing changes?

Question
Can consecutive 10-K differences be categorized, linked to exact evidence, and checked against earlier disclosures?
What was tested
150 issuer-disjoint companies, 150 pairs, 995 retained changes, and eight outcome-independent cases.
Result
The evidence workflow operated reliably. The bounded risk comparison was inconclusive.
Decision
Supported product Advance to an analyst workflow pilot—not trading deployment.
Lesson
A product can reduce research friction even when a predictive market claim is not established.
Locked product counts

629 high-confidence changes; 1,990 verified excerpt offsets; 394 previously disclosed; 342 genuinely new relative to the bounded study sources; 243 partially anticipated; 16 unclear; zero timing violations. The novelty counts sum to 995.

Not claimed: proven alpha, causal or validated risk prediction, customer adoption, willingness to pay, or institutional validation.

Next evidence gate

Test the analyst workflow.

Measure speed, recall, correction burden, trust, repeat use, and stored-context value.