Case study 10 / 11
AI voice and customer service
LiveClient work · NDAAI Lead Discovery on Reddit, With a Person Approving Every Reply
I was the sole engineer on an AI system that finds small business owners on Reddit describing problems an AI phone agent could solve. It judges each post with AI, drafts a reply in the founder's voice and sends it to Slack for approval. The system cannot post to Reddit at all. A person decides every reply.
System
Read 56 Reddit communities
System
Filter for real prospects
AI
AI judges intent and drafts a reply
Person
Approve, edit or skip in Slack
Person
A person posts it
On this page
The challenge
The founder of an AI phone agent company was finding customers by hand. He scrolled Reddit for small business owners complaining about missed calls or a chaotic front desk, wrote a genuinely helpful reply, and sometimes turned that into a conversation. It worked, but it depended on him being in the right community at the right hour.
He wanted it to scale without turning into spam. Reddit communities punish self-promotion quickly and permanently, and his personal account was the asset at risk. So the brief had one hard rule from day one: the system finds and drafts, and a person decides and posts.
What I delivered
A system that reads the right communities
The system checks 56 Reddit communities across 14 small business sectors, including trades, dental, beauty, hospitality and professional services, every four hours. Paid commercial access to Reddit was far beyond the budget at this stage, so it reads Reddit's public community feeds, which need no special account. I measured how those feeds really behave rather than trusting the documentation, and I also researched the paid options and their prices so the client could decide on real numbers.
A filter that finds real prospects
Posts pass through several checks, cheapest first, so the AI only reads posts worth paying for. A post is a candidate only when it comes from a business owner in a target sector and describes a problem the product can help with. A plumber talking about anything is not a lead. Anyone talking about missed calls is not a lead either. The overlap is. Mentions of competing tools count as the strongest signal, because someone comparing tools is already looking to buy.
Telling a business owner from everyone else was harder than it sounds. An early version rejected the best lead in the database because the post said "our home care agency" and the system only knew phrases like "we're a". I replaced the phrase list with pattern matching, so "my cleaning company" and "our agency" both count while "my wife's salon" does not.
A review queue in Slack
Every qualifying post becomes one card in Slack. The founder's first feedback was that the card mixed the story and the reply, and he was right: reading what someone said and judging what to say back are two different jobs. I rebuilt the card with the original post and the proposed reply in two clearly separated colours, plus a label that shows at a glance which kind of opportunity it is:
- Likely buyer: someone with a phone or call handling problem.
- Worth a helpful reply: an owner with a different business problem, where the value is being useful first and known second.
The founder can Approve, Edit, Rewrite or Skip. Only named people can approve, and a rewrite replaces the card instead of stacking new ones, so the channel stays at one card per post.
I also added a daily summary of what the filter turned down, with a Should have reached me and a Fair call button on each one. Until then the founder had no way to disagree with the filter, so it had no way to learn from him.
Drafts that learn the founder's voice
- When he edits a draft before approving it, the before and after are saved and used as examples for future drafts.
- When he gives the same instruction three times, such as "make it shorter", it becomes a standing rule.
- When he pastes a finished reply instead of an instruction, the system recognises it as an edit. Without that, a whole pasted comment would have become a permanent rule applied to every future draft.
How it works
- Read: public feeds from each community are collected, gently paced to respect Reddit's limits.
- Screen: old posts, empty posts and bot posts are removed.
- Match: the post must come from a target sector and mention a relevant problem.
- Judge: AI assesses intent and returns a structured verdict.
- Draft: AI writes a reply in the founder's voice, checked against length, tone and content rules.
- Review: the card arrives in Slack. A person approves, edits or skips it and posts by hand.
Key decisions
- No ability to post, at all. It isn't disabled or hidden behind a setting. The feature does not exist. That was a product decision to protect the founder's reputation, not a technical limit.
- Measure before building more. Qualified leads were arriving slowly, and I assumed the filter was too strict, since it turns down about 98 percent of what it reads. Before spending a week on a more complex approach, I had AI re-read 550 rejected posts from scratch without knowing they had been rejected. It disagreed with none of them. The filter was not the problem, and that test saved the client a week of unnecessary work.
- Build the outcome, not just the specification. The client had given me two documents: a detailed filter specification and a broader brief about being a useful peer to small business owners. I had built the first as if it were the whole job. The brief included three example posts of ideal prospects, and when I tested them the system rejected all three. One, a solo plumber deciding whether to hire his first employee, was blocked because "hiring" was on a list I had added to screen out job adverts. I widened the product to cover any operational problem a knowledgeable peer could genuinely help with, across staffing, scheduling, customer communication, quotes and tools.
- Check both directions. After widening the filter I measured quality again and found that about one in seven new catches was noise, such as software founders talking shop. I caught that from the data, before the client lost confidence in his queue.
Built for trust
- Reputation safeguards. At most two comments per community per week, no mentions of competitors in any draft, and every draft checked for length, tone and formatting before a person sees it.
- Nothing silent. Almost every serious issue in a background system like this has the same shape: the run reports success but quietly does less than it claims. I designed against that. A second scan can't run at the same time as the first, a busy community is retried rather than dropped, one bad post can't stop a whole sweep, and anything deliberately held back is logged with a reason.
- Verified sources. Every community is checked against Reddit before use, because Reddit quietly redirects unknown names instead of reporting an error. Twelve were rejected, each with a recorded reason.
- Visible data decisions. Raw posts are normally kept for 48 hours. That limit is paused on purpose while the filter is being tuned, and the system states it on every run so the decision never becomes a habit.
- Secure by design. The server accepts no incoming connections, passwords live only in AWS's secure parameter store, and deployments use no stored cloud keys.
The results
- 56 communities across 14 sectors checked every four hours, with every community covered in each full sweep.
- Close to 6,000 posts assessed and more than 40 qualified prospects surfaced, including a home care agency evaluating phone systems, a catering business choosing an AI phone agent, and a dispatch company scaling past a hundred contractors.
- An eightfold increase in qualified prospects after correcting the scope against the client's brief, with sectors that had never produced a lead starting to produce them. One was a barber looking for an affordable way to handle booking calls while cutting hair.
- 474 automated tests running against a real database, plus 49 real user journeys checked against the live system.
- Clear client reporting: a full filter report explaining every rule and rejection, a slide deck for review calls, and a project board written for a non-technical reader.
Under the hoodShow technical detailsHide technical details
- Access: Reddit's public per-community RSS feeds. A custom User-Agent header is required, since without it Reddit returns 403, which looks like bot blocking but isn't. The anonymous rate limit is about one request per 30 seconds. When Reddit tightened it to around 57 seconds mid-project, I caught the change from response headers. Each feed returns only the newest 25 posts, so I measured real coverage: 82 percent of all posts seen, with 54 of 56 communities at 100 percent.
- Pipeline: fetch, structural screen, topic match (vertical context AND at least one trigger topic, plus 61 tracked competitor and platform names), LLM intent judgement with a structured schema, then drafting with a policy check and automatic revision when a draft fails.
- Slack: Socket Mode, so the bot dials out and needs no public webhook. Block Kit cards use colour-separated attachments. Approvals are restricted to named users, and rewrites update the card in place.
- Infrastructure: a small ARM-based EC2 instance running three containers (PostgreSQL, scheduler, Slack bot). The security group has zero inbound ports, and administration is done entirely through AWS Systems Manager. Secrets are pulled from Parameter Store by the instance role at deploy time, and the database URL is assembled on the server so the password exists in one place.
- Deployments: GitHub Actions with OIDC federation into a deploy role and command dispatch through Systems Manager, so no long-lived AWS keys exist anywhere. The new build only replaces the running system if it succeeds. I also fixed a subtle Docker Compose issue where images were named after the build folder, so containers reported healthy while still running the previous version. The project name is now pinned, and each deploy compares the source on disk with the source inside the running container.
- Reliability: a cross-process lock so concurrent scans stand down and say so, separate exception types for "not yet" (rate limited) and "never" (missing community), a catch-up pass for posts judged but never carded, and per-post error containment.
- Testing: 474 tests run in CI against a real PostgreSQL service, not only SQLite, after a column length that SQLite ignores broke one button in production. Regression tests use the exact wording of the real posts that caused each bug, CI fails if any test is skipped, and a syntax-level test keeps the entry point at the end of each module.
What I'd do differently
- Read every source document against the others on day one. I built the specification I was given rather than the outcome described in the brief, and the clue was in the brief's worked examples all along.
- Measure missed opportunities before tuning quality. Every review looked at the posts that got through, because those are visible. Nobody looked at the rejected pile until I counted it.
- Treat "it reported success" as a claim to test. Half the serious bugs on this project were invisible because the part that failed also wrote the report.
Related work
- Call Agent AILive
Universal CRM Connector for an AI Calling Agent
Instead of writing a separate integration for every CRM, I built one universal connector that lets an AI calling agent work with many CRMs, and delivered a full Jobber integration end to end.
AI voice and customer service
- CS2 Technologies productLive
GWS Connect 24: Adding AI to a Live B2B Marketplace
AI added to a live wholesale marketplace: a 24/7 AI sales agent, deals sent to the CRM in one click with their activity shown back, and AI that drafts personalised outreach for a person to approve.
B2B wholesale commerce
- CS2 Technologies productIn production
AgentlyLeads: An AI-Native CRM You Run from ChatGPT and Claude
A CRM that teams operate straight from the AI assistant they already use. The AI prepares emails and updates, and a person approves anything that is sent, deleted or charged.
Sales and CRM