← Projects

WhatsApp automation for a CRM, with AI doing the classification

A Python script that reads conversations through the Evolution API, classifies them with the Anthropic API and feeds the CRM on its own every morning, with database and processing on the machine itself.

PythonAnthropic APIEvolution APISQLiteFlask
Status
In production
Year
2026
Context
Own product, in daily use
07:30
Runs on its own every day
Local
Database and processing on the machine
WhatsApp automation for a CRM, with AI doing the classification

The problem

All client contact happened on WhatsApp, and the history lived only there. Knowing who asked about what, in which month, and what was agreed meant scrolling through hundreds of conversations. The data existed, it just was not queryable.

What I built

  • Collection. The Evolution API exposes the conversations; a Python script reads the messages that are new since the last successful run.
  • Classification. Each conversation goes through the Anthropic API, which returns the funnel stage and a structured summary. For a contact that already exists, the current CRM state goes into the prompt, so the model consolidates instead of overwriting whatever the latest conversation window does not mention.
  • Persistence. A local database, with per-message deduplication so that re-runs do not inflate the history.
  • Scheduling. Windows Task Scheduler fires at 07:30, with recovery for when the machine spent that time switched off.
  • Triage queue. A separate sweep analyses the older history, which the daily sync window never reached, and drops the result into a queue for review.
  • Catalogue mirror. It consumes the public API of the website and keeps a local cache, so the contact record already shows the item the person asked about.

Technical decisions

A simplification that also solved privacy. The first version went through a cloud workflow orchestrator. It worked, but it was one more piece to maintain and it sent client conversations over the wire for no reason. Replacing it with a scheduled script made the system both simpler and more private at the same time, which is the rare combination. The orchestrator was retired from this flow.

AI output treated as untrusted input. The classification comes from a model, so the response is validated against a closed set of stages before it touches the database. If something unexpected comes back, the conversation is left unclassified instead of storing an invented value. I prefer an empty field to wrong data that nobody will audit later. Along the same line, there are guards in code and not only in the prompt: the model never closes a deal on its own and never moves anyone out of their own category.

The human note beats the machine reading. Notes live in their own table, append-only and dated. The AI reads those notes with precedence over the conversation and never rewrites them. Before that, they shared a field with the automatic summary and disappeared on the next analysis, with no error at all. They are two sources with different owners.

The sweep does not write to the funnel. Analysing the older history could scramble months of classification in one go, so it writes to a separate queue and a single click promotes someone into the base. Cost was not the obstacle, curation was.

Deduplication at the source. Re-running is normal, whether from a network failure or a machine left off. Without a per-message deduplication key, every re-run would duplicate the history.

How I verified it

Re-running over the same window. The first real load created 23 contacts, updated 1 and threw no errors; the second pass, over the same interval, was 100% deduplication.

A pilot before turning the sweep loose on the whole history. I ran it on 15 conversations and checked them one by one: 13 were legitimate contacts. The hard case was the one that interested me most, because the longest conversation in the queue was from a supplier and the classifier got it right. A filter by conversation length, which was the obvious and cheap route, would have put that very one at the top.

Testing the connection guard in both directions. This is where the most dangerous problem in the project was: with WhatsApp disconnected, the API kept responding normally, serving from its own cache. The sync would run, find no new messages and record “0 created, 0 updated, 0 errors” for weeks. Apparent success is worse than an error, because nobody is going to investigate it. I started checking the connection state before anything else and aborting when it is not open, treating “I could not verify” as a failure too. I tested it disconnected, and it aborted; I reconnected, and it imported 62 conversations.

It was the third case of the same family in this project. Before it, a backup routine that failed returning only an error code nobody read, and a misconfigured username that inverted conversation metrics without ever throwing an error. Since then, every scheduled routine writes the result of its own run to a file, so that silence stops looking like success.