Investabot

An engineering case study: what it is, what failed, and what was rebuilt.

Investabot's Market Insights feed showing automated research articles generated from SEC filings
Project
Investabot
Role
Founder & Developer
Started
November 2025
Area
Financial technology, financial research software
URL
investabot.ai
  1. 01Overview
  2. 02Why I built it
  3. 03How it evolved
  4. 04Architecture
  5. 05The SEC data pipeline
  6. 06Financial table extraction
  7. 07Validation and reliability
  8. 08AI's role
  9. 09Problems I encountered
  10. 10Failed approaches
  11. 11What I rebuilt
  12. 12Current development
  13. 13What comes next
01Overview

Investabot is an automated financial research platform built around primary-source SEC filings. It monitors disclosures (8-K, Form 144, Form 3, 13F) across the S&P 500, extracts financial facts from the underlying documents, and generates sourced research articles with validation and review gates in front of publication.

I founded it in November 2025 and have built and operated every part of it alone.

02Why I built it

I built Investabot around SEC filings because I wanted the research to remain tied to primary-source financial information. Filings are public, authoritative, and structured enough for software to work with, which makes them the best possible material for the intersection I care about. The goal was simple to state: research where every number is traceable to the disclosure it came from. Building a system that actually meets that bar turned out to be the entire project.

03How it evolved

Investabot did not begin as what it is now. Early versions centered on deterministic company scoring and educational tooling. Over time the center of gravity moved to generating research articles directly from filings, and that shift raised the reliability bar dramatically: a scoring engine can tolerate a rough input, but a published article cannot tolerate a wrong number. Most of the platform's later engineering follows from taking that seriously.

04Architecture

The platform is a Node.js and Express backend with PostgreSQL and Prisma, and a React frontend. Conceptually it is a pipeline: filing ingestion, exhibit resolution, financial data extraction, validation, AI-assisted article generation, and review before anything is treated as publishable. Each stage is designed so the next stage can trust its output, and the pipeline prefers producing nothing over producing something unverified.

05The SEC data pipeline

The system watches EDGAR for new filings from covered companies. A filing is usually a bundle of documents, and the financial content is rarely the first file: for an 8-K earnings release, the press release typically lives in Exhibit 99.1, so the pipeline resolves exhibits to find the document that actually contains the numbers. Where XBRL structured data is available, it is used as a stronger source than presentation HTML.

06Financial table extraction

This is the hardest problem in the system. Earnings tables in the wild use multi-row headers with colspans, so the period and the year for a column can live in different rows and cells. Before reading a single value, the parser reconstructs the effective header of every column.

  • Period identification. Quarterly, year-to-date, and prior-year columns sit side by side. Attributing a value to the wrong period produces a plausible wrong number.
  • Scale and units. "In thousands" versus "in millions" is stated once, far from the values it governs, and a scale error still produces a number shaped like a real one.
  • GAAP versus adjusted. Companies present both, often interleaved. The pipeline separates them explicitly rather than taking whichever appears first.
  • Concept boundaries. Consolidated versus attributable to shareholders, revenue versus deferred revenue: adjacent concepts that must not contaminate each other.
07Validation and reliability

Extraction is tested against a corpus of human-verified expected values taken from real filings. Every change to extraction logic is diffed against that corpus before it ships, and validation gates sit between extraction and anything user-facing.

Plausible but wrong financial information is more dangerous than missing information. The system abstains when evidence is insufficient rather than inventing or mislabeling a financial fact.
08AI's role

AI assists with article generation, not with deciding what the numbers are. Generation is constrained to facts that deterministic extraction has produced and validation has accepted, and drafts pass through review requirements rather than publishing unattended. Model output is treated as untrusted until it is checked against the extracted data.

09Problems I encountered

The defining problem was silent failure. Early extraction failures were not obvious garbage: they were plausible numbers that were wrong. Values off by a factor of a million because of a scale mistake. Figures pulled from the wrong period column. Adjacent concepts mixed together. A few of those wrong numbers reached published articles before I understood the failure class; they were corrected, and that event reshaped the project.

10Failed approaches

The first extraction approach worked directly over the text of filing HTML with pattern matching. It worked often enough to be dangerous and failed in exactly the ways described above.

My first response to those failures also failed: patching individual bugs as they appeared. Each fix was real, but fixes tuned to one company's format sometimes regressed another's, and I kept forming theories about failures that turned out to be wrong when tested against the actual documents. Debugging by theory, without ground truth to check against, wasted more time than any single bug.

11What I rebuilt

The turning point was to stop theorizing and build ground truth. I put a mandatory review gate in front of the legacy pipeline so nothing more could publish unattended, then built an offline test harness around human-verified expected values from real filings, so any extraction change could be checked against known-correct answers instead of my assumptions.

Then I rebuilt extraction itself as a deterministic, structure-aware system rather than a pile of patterns, using XBRL structured data where it is available. The new pipeline runs in shadow alongside the legacy one on live filings, and every disagreement between the two gets investigated and attributed before the new system earns trust. So far, investigating those disagreements has consistently found the error on the legacy side, which is exactly the evidence a cutover decision should be built on.

12Current development

Current work is the tail end of that migration: growing the verified test corpus, running the shadow comparison, and holding the cutover to explicit, written criteria instead of a feeling that it is probably fine. Alongside it, I am building out the platform's public-facing surface.

13What comes next

The goal is a system that reliably transforms primary-source corporate disclosures into structured, traceable financial information and research: not the most content, the most trustworthy content per filing. The engineering standard I hold it to is simple: every published number should survive being checked against the document it came from.

Visit investabot.ai