Present
←
1 / 12
→
5 min
10 min
Notes
Jump to demo
Print
01 The customer problem
Institutional research workflow
Analyst attention—not information—is the scarce resource.
Filings, alerts, headlines, and internal notes arrive continuously. Analysts still have to decide what deserves review, rebuild history, and defend the conclusion.
Too much information Manual comparison Repeated searching Fragmented memory
Innovation frame
Target user Institutional research analyst
Buyer Director or Head of Research
Gatekeeper CTO or Head of Data
Beachhead Changed SEC disclosures
Speaking prompt · 20 sec Open with the queue, not the technology. The problem is the cost of deciding what to read, reconstructing what changed, and supporting that judgment when the team asks why.
Next Current tools organize information—but they do not resolve the attention decision.
02 Current tools
The opportunity is not another feed. It is a governed layer between incoming information and analyst judgment.
Product wedge: rank changed disclosures, show the historical comparison, and keep the source evidence attached.
Speaking prompt · 20 sec Existing systems are inputs, not competitors to replace. Pure News Intelligence complements them by resolving the comparison and provenance work between the alert and the analyst’s conclusion.
Next Six interviews shifted the design from summary generation to evidence-backed prioritization.
03 Customer discovery
N = 6 completed interviews
Trust, fit, and prioritization mattered more than another summary.
Paraphrased findings only. Titles, employers, and direct quotations are not inferred.
Interview perspective, workflow need, trust requirement, and product implication
Interviewee · perspective Workflow problem Trust requirement Product implication
Matthew G. SurprenantSearch, Q&A, usability Retrieval friction Answers trace to evidence Source-linked excerpts
Bilal MustafaSystematic, cross-asset Noisy prediction target Claims match evidence Separate product from alpha
Matt DiCensoFundamental workflow Manual prioritization Historical context Rank changes, show prior text
Joseph MarksInnovation framing Problem too broad Testable value Narrow the beachhead
Redmond CostelloEnterprise implementation Workflow disruption Integration and security Complement existing systems
Krishna ValluruData, governance, analytics Fragile research memory Provenance and control Governed evidence record
Users needed Prioritization · comparison · visible evidence · relevance · low friction
Buyers and gatekeepers needed Productivity · adoption · integration · security · auditability · governance
The product changed Evidence-first ranking · side-by-side history · institutional memory · analyst accountability
Still unvalidated: measured time savings, repeat use, willingness to pay, buyer approval, and enterprise integration.
Speaking prompt · 25 sec The interviews did not validate demand. They exposed design requirements: prioritization over generic summarization, visible evidence, workflow fit, and analyst accountability. Adoption and willingness to pay remain open questions.
Next The interviews challenged a broad trading hypothesis that the empirical work had already put under pressure.
04 The original hypothesis
The first experiment
Could news and filing language add incremental investment signal?
The original project treated document information as a potential input to return and risk prediction.
News + filings Point-in-time documents
→
Language features Model representations
→
Incremental signal? Locked empirical test
→
Investment outcome Return or downside risk
Claim under test
Added documents improve out-of-sample prediction.
Evidence standard
Locked comparisons, issuer-disjoint samples, and point-in-time inputs.
Decision rule
Do not turn a null or inconclusive result into a product claim.
Speaking prompt · 20 sec This was a legitimate hypothesis, not a promise. The important design choice was to set an evidence standard before interpreting results and to keep the investment claim separate from the document-processing capability.
Next The locked tests did not support the alpha claim—and the project preserved that result.
05 The negative results
Disciplined innovation
The market claims were not supported. The evidence-processing capability was.
Negative and null results were preserved rather than tuned away.
Not supported Incremental news or 10-K alpha Adding 10-K information did not improve the existing 8-K and inexpensive-news model.
Not supported Better expected-language reconstruction The prior 10-K alone predicted the next filing’s representation better than augmented models.
Inconclusive Disclosure-change downside risk The product found evidence-backed changes; the bounded risk test did not establish a reliable signal.
Final research classification INSUFFICIENT_EVIDENCE
Speaking prompt · 25 sec Say this plainly: the research did not prove alpha, causal prediction, or a validated risk signal. It did prove the technical ability to compare disclosures with point-in-time evidence. That distinction created the pivot.
Next The null result became a design input: refine the problem, not the statistics.
06 The product pivot
Evidence changed the project
The research did not prove alpha. It revealed a better innovation opportunity.
Hypothesis tested. Result rejected. User problem refined. Solution redesigned. Next experiment defined.
Original idea Predict returns
→
Empirical test Lock the comparison
→
Null result Reject the claim
→
Customer discovery Listen for workflow pain
→
New problem Focus analyst attention
→
Working product Evidence-linked changes
→
Analyst pilot Test workflow value
Evidence changed the project rather than being tuned away.
Speaking prompt · 20 sec Frame the null result as a decision gate. It prevented an unsupported trading pitch and redirected the project toward the concrete job interviews revealed: deciding what to read and preserving why it mattered.
Next The redesigned workflow turns a filing into a ranked, evidence-linked research record.
07 How the workflow works
1 Ingest Preserve source and filing time
2 Compare Match prior and current sections
3 Classify Category, direction, materiality
4 Check earlier disclosure 8-Ks and pre-filing headlines
5 Explain Why it may merit review
6 Preserve Export an auditable record
KHC The Kraft Heinz Company
High materiality Previously disclosed Confidence 0.947
Debt and refinancing language became more specific.
Added: acceleration, covenant breaches, waiver dependency, default, and access to the Senior Credit Facility.
Speaking prompt · 25 sec Show the handoff from system to analyst. The system narrows the queue and attaches evidence. It does not decide whether Kraft Heinz is an investment or predict what the stock will do.
Next One frozen Kraft Heinz case shows how comparison and novelty work together.
08 One real case
Kraft Heinz · 2023 10-K · Item 1A
A more specific debt warning—but not a wholly new issue.
Selected for evidence diversity and clarity, not later market performance.
Available at filing time Prior 10-K + current 10-K + strictly earlier 8-K and headline evidence
Prior filing
General refinancing risk “Changes in financial and capital markets … may increase the cost of financing as well as the risks of refinancing maturing debt…”
Current filing
Specific covenant consequences “The creditors who hold our debt could accelerate amounts due … [and we could be] unable to access our Senior Credit Facility .”
What changed Default, acceleration, waivers, and facility access were added.
Novelty Previously disclosed An intervening 8-K and earlier headline existed.
Analyst judgment Reconcile covenant headroom and waiver status.
Speaking prompt · 30 sec Read only the highlighted phrases. The new 10-K language is more specific, but the topic appeared earlier. That distinction prevents the analyst from confusing sharper wording with a first disclosure.
Next The product demonstration is supported; alpha, adoption, and workflow benefit are not.
09 Experimental extension
From evidence to a research question
A verified change can become a falsifiable research agenda—after a full-prior check. RH demonstrates a grounded hypothesis with explicit uncertainty, not a recommendation.
Verified change
→ Full prior-evidence check
→ Economic mechanism
→ Falsifiable hypothesis
→ Skeptical review
→ Analyst decision
RH · PASS_HYPOTHESIS_ONLY Evidence coverage is 1.00. The result is a research question requiring analyst and market-data checks.
DVN · PASS_HYPOTHESIS_ONLY A second bounded case passed as a hypothesis only. Neither case produced an actionable trade view.
Human review pending Independent evidence review and automated grounding checks are complete; formal blinded human ratings remain pending.
Open the frozen RH hypothesis · Open the frozen DVN hypothesis
Speaking prompt · 30 sec The engine creates research agendas, not predictions or recommendations. It generated two bounded, falsifiable hypotheses, both explicitly hypothesis-only. Formal human ratings remain pending, and no improvement in Sharpe ratio or information coefficient has been demonstrated.
Next A safe rejection can be as valuable as a publishable research question.
10 Safe failure matters
EFX grounding correction
Full-prior retrieval stopped a plausible but unsupported story. Safe rejection is a successful system outcome.
Missing context found The selected prior excerpt omitted the already-disclosed $125M conditional top-up.
Roles corrected $346.7M was a remaining balance; approximately $345M was a cash deposit. The quantities were not substitutes.
Mechanism pruned Unsupported liquidity, reserve, and bondholder claims were removed.
Safe disposition No actionable trade view was produced.
Two cases produced grounded research hypotheses. Six were safely rejected or held. None produced an actionable trade view.
View the frozen EFX grounding rejection
Speaking prompt · 25 sec The first version made an existing term look new. V2 searched the full prior record, corrected the financial roles, and removed the unsupported causal chain. That refusal is product value. Human ratings remain pending.
Next The evidence boundary remains explicit.
11 Evidence boundary
Supported versus unsupported
A working product demonstration is not the same as a validated business or investment result.
Four labels keep the claims legible.
Supported Reliable document processing Historical comparison Evidence-linked change detection Point-in-time integrity Reproducible classifications
Not supported Incremental news alpha Improved expected-language model Causal prediction Customer adoption Willingness to pay
Inconclusive Downside-risk signal Economic usefulness of risk labels
Not tested Analyst time savings Recall and trust Repeat use Institutional-memory value
Speaking prompt · 20 sec The site separates capability, research outcome, and commercial validation. The only green-light claim is that the bounded evidence workflow operates as demonstrated. User benefit still needs a pilot.
Next The next experiment is a controlled workflow pilot—not another unsupported trading claim.
12 Pilot and discussion
Next experiment · 2–4 weeks
Can 3–6 analysts reach better-supported conclusions faster?
Compare the normal workflow with Pure News Intelligence-assisted review on a bounded set of 8-K and 10-K changes.
Primary measure Time to identify and explain the most important change
Quality guardrails Recall, agreement, correction burden, and evidence quality
Adoption signals Repeat use, exports, stored-context usefulness, and preference
Questions for the professor
Is the customer problem focused enough? Which user or buyer is still missing? Which pilot metric should be primary? What evidence would justify a go decision?
Speaking prompt · 30 sec End with a decision, not a pitch. Ask which metric should govern the pilot and what evidence would warrant continuing, revising, or stopping. The requested help is feedback, an analyst introduction, or enterprise-requirement critique.
End Return to the overview or open the live workflow.