Posts  / #POST-221920
REDDIT

What insights would you try to extract from a large dataset of SEC comment letters?

P
Mar 14, 2026 · 01:50

I’ve been working on aggregating SEC comment letters and company responses from EDGAR into a dataset so they’re easier to analyze. The filings are public, but they’re scattered and not particularly easy to explore systematically.

When I first started digging into the data, I expected there might be some obvious patterns like certain sectors getting far more scrutiny than others. But at a high level it mostly seems to correlate with the number of companies in each sector. Bigger sectors naturally generate more correspondence.

That said, I’m pretty confident there are still meaningful insights buried in the data they’re just probably not visible from simple counts.

One direction I’m thinking about exploring is analyzing the actual text of the letters to see if the SEC starts asking similar disclosure questions across multiple companies in the same industry around the same time. If that happens, it could potentially reveal areas where regulators are beginning to focus before it becomes widely discussed, things like accounting treatment, metrics companies report, or disclosure practices that might later force companies to revise filings.

Curious how others would approach this. If you had a large dataset of SEC comment letters, what signal or insight would you try to find?