← Projects

Qualifying a lead base from open data

Brazil's public company registry processed in Python until it became a useful list. Along the way, the discovery that the data did not serve the use I had imagined for it.

PythonPandasCNPJ Open Data (Brazil)
Status
Pilot
Year
2026
Context
Own product
1,292
Records after filtering
Public
Source of the data
Qualifying a lead base from open data

The problem

Deciding who to offer a product to without buying a ready-made list, which is expensive and of poor quality.

What I built

Brazil’s federal tax authority publishes the complete company registry (CNPJ), in large, denormalised files. The work was processing it in Python until it became something queryable:

  • A geographic cut for the region served.
  • A company-size filter, excluding sole-trader micro-entrepreneurs: the product’s price point does not match the financial reality of that profile, so including them would only inflate the list.
  • A filter by economic activity, restricted to the lines of business where the product makes a difference.

Result: 1,292 records with a compatible profile, out of millions of rows.

The discovery that changed the conclusion

Checking a sample against reality turned up the problem: the phone number and email registered in the database are usually the accountant’s, not the business owner’s. The data is public and true, it just is not the channel I assumed it was.

That invalidated the original use (reaching out directly) and changed the conclusion: the list is good for knowing who exists and where, supporting an in-person or referral-based approach. I recorded the limitation instead of carrying on with the wrong premise, which would have cost weeks of unanswered outreach and a mistaken reading of “the market does not want this”.

It is the most direct example of how I work: check the output against reality before acting on it. The code was right the whole time; it was the premise that was wrong, and no unit test would have caught that.