Proposed next experiment · 2–4 weeks

Test workflow value before making a commercial claim.

Can 3–6 analysts identify and explain the most important disclosure change faster without reducing evidence quality?

Pilot design

A bounded side-by-side workflow comparison

The analyst remains responsible for the final conclusion.

Participants

3–6 analysts

Enough for observed workflow variation without pretending to establish broad adoption.

Task

Review 8-K and 10-K disclosure changes

Use a bounded issuer and filing set with known source availability.

Comparison

Normal vs assisted workflow

Compare current process with Pure News Intelligence-assisted review.

Duration

2–4 weeks

Long enough to observe repeat use, corrections, and research-record value.

Example pilot scorecard

Proposed measurement—not actual performance

Baseline and assisted values remain blank until a controlled pilot occurs.

Proposed pilot measurements and evidence captured
MeasureRoleNormal workflowAssisted workflowEvidence captured
Time to identify and explain the most important changePrimaryTo measureTo measureTask timestamps + analyst explanation
Material-event recallSecondaryTo measureTo measureReference-set comparison
Category and materiality agreementSecondaryTo measureTo measureAnalyst labels vs system labels
Excerpt correction rateGuardrailNot applicableTo measureCorrections per reviewed record
Novelty-label acceptanceTrustNot applicableTo measureAccept / revise + reason
Repeat use and exportsAdoption signalTo measureTo measureSessions + evidence packets
Stored-context usefulness and preferenceQualitativeTo measureTo measurePost-task interview

Predefined decision rules

Go, revise, or stop

Threshold values should be finalized before the pilot begins; these are directional criteria.

Go
  • Meaningful time savings
  • No reduction in evidence quality or recall
  • Strong repeat use and analyst trust
Revise
  • Time savings but weak label acceptance
  • Useful comparisons but excessive correction burden
  • Value concentrated in only part of the workflow
Stop
  • No meaningful time savings
  • Lower material-event recall
  • Low trust or workflow friction exceeds benefit

Help test the next version

Choose the smallest useful contribution.

No personal information is collected on this site.

Proposed idea-engine evaluation

Measure research-agenda quality before market performance

No results are reported here.

Workflow outcomes

  • Analyst acceptance rate
  • Accept-with-edit rate
  • Rejection rate
  • Time to create a research agenda
  • Edit burden
  • Analyst preference

Grounding quality

  • Grounding score
  • Novelty accuracy
  • Mechanism clarity
  • Falsifiability
  • Usefulness and non-obviousness
  • Affected-entity accuracy
  • Safe-rejection quality

Comparison groups

  1. Disclosure change only
  2. Generic AI summary
  3. Grounded AI hypothesis
  4. Grounded hypothesis plus skeptic
  5. Analyst-edited hypothesis
  6. Human-created research agenda

Sequencing rule: Only after hypotheses are frozen prospectively should a separate study evaluate IC, Sharpe ratio, hit rate, signal decay, or realized-versus-implied volatility.