Honestly, I wrote this article months ago. I wish I could simply say I have been busy, but that would be an understatement.
First, Click Here Digital recently passed its SOC 2 Type II audit. That initiative required significantly more work than I anticipated, and I now fully understand why organizations have dedicated teams for herding those cats. I usually aim to avoid spending too much time in HR offices, so that process alone gave me anxiety. Additionally, I have been focused on a new passion project after clocking out, which I will share more about in a future article.
But the primary reason for the delay was MADI, Click Here Digital's branded AI agent. While it is currently an internal-only tool, its potential makes it a game-changer for our operations. I will not go into the specific details of its capabilities to protect proprietary information, but I do want to highlight some of the architectural design decisions I made and the technical lessons I learned during the process.
I built MADI's first iteration in partnership with IntellegenAI, a generative AI consulting business owned by Gilberto Mizrahi. We all had ideas about what could be possible, but we did not know what was realistic or where to start. Gilberto and I worked together to design the initial workflow using our existing APIs as a proof of concept for myself and the leadership team—he brought the generative AI expertise, and I brought the deep knowledge of our data, systems, and business context. Out of that collaboration came a blueprint detailing the tools, how to use them, and what to monitor. Then, it clicked. I began experimenting. I built workflows, I broke things, and I fixed them—sometimes. I became obsessed with the potential.
However, through this process, I soon realized MADI's biggest risk. Marketing data demands absolute precision. If a client asks an account manager for last month's ad spend, "$125,450" is the only acceptable answer. "$125k-ish" or a number hallucinated by a model that got confused by a column name is not just a wrong answer; in my line of work, it is a liability.
The solution to building a reliable enterprise tool was not to force the LLM to do the math. The solution was to stop treating the LLM as the worker and start treating it as the router.
What I landed on was an Intent-Driven Workflow, which I designed to combine the flexibility of natural language with the deterministic precision of code.
The "Chat-to-SQL" Fallacy
The standard approach to AI analytics usually looks like this: You feed a database schema to an LLM, the user asks a question, and the LLM tries to write a SQL query to fetch the answer. This is the "Hello World" of AI engineering. It is impressive in a demo, but it falls apart in production.
Real-world business questions are rarely simple. A query like "How are we doing compared to last month?" requires context. It requires pulling data from multiple disparate sources—Google Ads, Facebook, Bing, Programmatic Display, and CTV. It requires complex date math to determine what "last month" actually means relative to today. It requires understanding that "performance" means different metrics for a video campaign versus a search campaign.
When you ask an LLM to handle all of that in one pass, you get drift. The model might hallucinate a metric, fail to join tables correctly, or simply timeout while trying to "think" through the logic.
I realized that to build a system that our account managers and executives could trust—one that could reduce a 3-hour reporting process to under 30 seconds—I needed to decouple the intent (what the user wants) from the execution (how we get the data).
The Solution: An Intent-Driven Architecture
I architected MADI around a three-stage pipeline that ensures safety, scalability, and, most importantly, accuracy. In this system, the LLM never touches the database directly. Instead, it functions as a translation layer, converting messy human language into a structured "intent"—a JSON blueprint that my code then executes deterministically.
Here is the high-level flow of how a user's question translates into a report:
┌─────────────────────────────────────────────────────────┐
│ USER INTERFACE │
│ "How is Acme Corp doing this month across all │
│ channels compared to last month?" │
└────────────────────┬────────────────────────────────────┘
│
│ Natural Language Query
↓
┌─────────────────────────────────────────────────────────┐
│ AI INTENT GENERATION ENGINE │
│ ┌────────────────────────────────────────────────┐ │
│ │ Step 1: Understand Intent │ │
│ │ • "Performance query" identified │ │
│ │ • Intent type: client_performance │ │
│ └────────────────────────────────────────────────┘ │
│ ┌────────────────────────────────────────────────┐ │
│ │ Step 2: Extract Time Periods │ │
│ │ • "this month" → 2025-12-01 to 2025-12-11 │ │
│ │ • "last month" → 2025-11-01 to 2025-11-30 │ │
│ └────────────────────────────────────────────────┘ │
│ ... (Client & Channel Logic) ... │
└────────────────────┬────────────────────────────────────┘
│
│ Structured Intents (14 queries)
↓
┌─────────────────────────────────────────────────────────┐
│ PARALLEL DATA PROCESSING ENGINE │
│ │
│ Batch 1 (10 parallel): Batch 2 (4 remaining): │
│ ┌──────┐ ┌──────┐ ┌───────┐ ┌──────┐ ┌─────┐ │
│ │Search│ │Social│ │Display│ │Audio │ │ SEO │ │
│ │Dec/Nov│ │Dec/Nov│ │Dec/Nov│ │Dec/Nov│ │Dec/Nov│ │
│ └──────┘ └──────┘ └───────┘ └──────┘ └─────┘ │
│ (Video and CTV omitted for visual brevity) │
│ │
│ All 7 channels processed across both time periods │
└────────────────────┬────────────────────────────────────┘
│
│ Raw Performance Data
↓
┌─────────────────────────────────────────────────────────┐
│ ANALYTICS & INSIGHTS ENGINE │
│ Metric Calculation & AI-Generated Recommendations │
└─────────────────────────────────────────────────────────┘
Figure 1: High-Level Architecture showing the flow from natural language to structured intent to parallel execution.
Stage 1: The Brain (Intent Generation)
When a user types, "How is Acme Corp performing this month?", the system doesn't immediately look for data. It sends the prompt to the AI model (Google Gemini) with a single, highly specific goal: structure the request.
This is where the engineering mindset kicks in. I do not trust the AI to guess. I implemented strict validation logic. The AI analyzes the natural language and extracts critical parameters into a JSON object:
- Intent Type: It identifies this as a
client_performancerequest. - Client Recognition: This is one of the hardest parts of the pipeline. Users make typos. They say "Acme" instead of "Acme Corp." I built a fuzzy matching algorithm that validates the extracted name against our client database, requiring a minimum confidence score of 60% to proceed. If the confidence is lower, the system halts and asks for clarification rather than guessing.
- Time Period: It converts vague phrases like "this month" or "YTD" into precise ISO-8601 date ranges (e.g.,
2025-12-01to2025-12-11). - Channel Expansion: It understands that "all channels" isn't a single database table. It expands this request into our 7 key marketing tactics: Search, Social, Display, Video, CTV, Audio, and SEO.
┌─────────────────────────────────────────────────┐
│ AI NATURAL LANGUAGE UNDERSTANDING │
├─────────────────────────────────────────────────┤
│ │
│ Intent Classification │
│ ├─ Is this a performance query? │
│ ├─ Is this a budget query? │
│ └─ Is this actionable? → YES │
│ │
│ Date Intelligence │
│ ├─ "this month" → MTD (month-to-date) │
│ ├─ "last quarter" → Q3 2025 │
│ └─ Auto-validates against data availability │
│ │
│ Client Recognition │
│ ├─ Extract: "Acme Corp" │
│ ├─ Database fuzzy matching │
│ └─ Threshold: 60% minimum confidence │
│ │
└─────────────────────────────────────────────────┘
Figure 2: The logic flow for converting vague language into structured parameters.
This stage is purely about understanding. If the AI is unsure, it stops. But if the intent is clear, we move from the probabilistic world of AI to the deterministic world of engineering.
Stage 2: The Muscle (Parallel Execution)
Once we have the structured intent, I no longer need the LLM. I need raw speed and computational accuracy.
In a traditional synchronous system, checking 7 channels for two different time periods (to calculate period-over-period growth) would require 14 sequential database queries. If each query takes 2 seconds, the user is waiting half a minute just for data retrieval.
I designed the execution engine using Go (Golang) to leverage its superior concurrency primitives. The system takes the "Intent Object" and spawns a set of parallel workers.
For a single query like "Acme Corp performance," the system might generate 14 distinct data requests. Because I utilized a containerized architecture with a dedicated data warehouse, these requests run concurrently.
14 Intents Generated
│
↓
┌───────────────────────────────────┐
│ BATCH PROCESSING CONTROLLER │
│ (10 parallel workers) │
└───────────────┬───────────────────┘
│
┌───────────────┼───────────────────┐
│ │ │
↓ ↓ ↓
┌────────┐ ┌────────┐ ┌────────┐
│Worker 1│ │Worker 2│ ... │Worker10│
│ Intent │ │ Intent │ │ Intent │
│ #1 │ │ #2 │ │ #10 │
└────────┘ └────────┘ └────────┘
Figure 3: Parallel processing architecture that reduces sequential query time by over 85%.
This parallelism is the "secret sauce" that allows MADI to deliver a comprehensive cross-channel report in under 30 seconds. It is not magic; it is just good engineering fundamentals applied to an AI workflow.
Stage 3: The Voice (Synthesis & Recommendation)
Raw data is useless without context. A spreadsheet with 14 tabs of metrics is just noise to a busy executive. This is where I bring the AI back into the loop.
In this final stage, the system passes the calculated metrics—not the raw database rows—back to the AI with a new prompt: "Given these precise numbers, what is the recommendation?"
This is where the LLM shines. It does not have to calculate the CPA (Cost Per Acquisition); my code has already done that to the exact penny. The AI just needs to look at the trend (e.g., CPA dropped 3.8%) and generate the narrative: "Search and Video are delivering the lowest CPA; recommend scaling investment."
This hybrid approach ensures that the numbers are always 100% accurate (because they came from code), while the explanation is natural and human-readable (because it came from the LLM).
Leveraging the Model Context Protocol (MCP) for Scale
In designing MADI, I did not want to build a walled garden. I wanted a system that could easily interface with other data systems and future AI models. This led me to adopt an architecture based on the Model Context Protocol (MCP)—an open standard for connecting AI assistants to external tools and data sources through a consistent client-server interface, so any capability can be exposed as a modular "tool" the model can invoke.
In the context of MADI, this means that the core logic of the application isn't a monolith. It is a collection of discrete tools: date_context, get_clients, generate_intents, and process_intents.
Why This Protocol Matters
By standardizing these interfaces, I achieved two critical engineering goals:
- Horizontal Scalability: The system can handle more complexity without becoming more fragile. If I need to add a new capability, I don't have to risk breaking the existing
client_performancelogic. I just plug in a new tool. - Model Independence: The intent structure acts as a contract. I can swap the underlying AI model (e.g., moving from Gemini 2.5 to Gemini 3.0) without rewriting a single line of application code. As long as the new model can output the standard intent JSON, the rest of the pipeline works perfectly.
This tool-based design effectively turns MADI into an operating system for marketing intelligence. It is not just answering questions; it is orchestrating a fleet of specialized tools to perform complex work.
Future-Proofing via Tool-Based Extensibility
I try to think ahead as much as possible, looking past the current sprint to design for the "what if" scenarios—building a foundation flexible enough to absorb future requirements without forcing a complete re-architecture. The intent-driven architecture allows for unlimited functional extensibility without modifying core systems.
Let's say next quarter our strategy team demands a "Competitor Analysis" feature. In a traditional BI platform, this would be a 3–6 month project involving ETL pipelines and dashboard redesigns. With MADI's tool-based extension pattern, the process is significantly faster.
┌───────────────────────────────────────────────────┐
│ INTENT TYPE REGISTRY │
│ (Defines all supported analysis types) │
├───────────────────────────────────────────────────┤
│ │
│ CURRENT: │
│ date_context → date_context │
│ client_lookup → get_clients │
│ client_search → search_clients │
│ client_performance → generate_intents │
│ performance_report → process_intents │
│ │
│ FUTURE: │
│ budget_optimization → analyze_budget │
│ performance_forecast → predict_performance │
│ anomaly_detection → detect_anomalies │
│ competitive_analysis → competitive_intel │
│ │
└───────────────────┬───────────────────────────────┘
│
↓
Each Tool Is Self-Contained:
┌──────────────────────────┐
│ 1. Input: Intent object │
│ 2. Data retrieval logic │
│ 3. Analysis/computation │
│ 4. Output: Results │
└──────────────────────────┘
Figure 4: The Extensible Tooling Architecture allowing for rapid feature deployment.
I simply register a new intent type (competitive_analysis) and map it to a new tool function that hits a third-party API like SEMrush. The workflow remains identical:
- AI Layer: Recognizes the user wants competitive data and outputs a
competitive_analysisintent. - Router: Sees the new intent and routes it to the
process_competitive_analysistool. - Execution: The tool fetches the data and returns structured results.
This modularity reduces the time-to-market for new features from months to weeks. It allows the system to grow organically. We can add Budget Forecasting, Anomaly Detection, or Attribution Modeling as separate modules that never interfere with one another. This is the difference between building a script and building a platform.
The Engineering Reality Check
There is a concept I often discuss called "vibe-coding"—the idea of letting AI handle the heavy lifting while you just direct the "vibes." I've written about how dangerous this can be without a validation framework. MADI is the antithesis of vibe-coding. It is a rigorously engineered system where AI is a component, not the captain.
I validated every step of this workflow. I watched the agents fail, hallucinate, and misinterpret instructions, and I wrote the guardrails to prevent those failures from reaching the user. The result is a tool that does not just look cool in a demo but actually drives business value. It saves our team nearly 3 hours per complex query. It allows for real-time decision-making during client calls. It transforms our analysts from data gatherers into data strategists.
Conclusion
This architecture has transformed MADI from a novelty into a critical business tool. It gives us the best of both worlds: the natural language interface that makes data accessible to everyone, and the engineering rigor that ensures that data is accurate.
For a deeper dive into how this powers our ClickIQ platform, you can read more at Click Here Digital Technology and see the MADI announcement.
True technical leadership is not just about using the newest models. It is about knowing when to let the AI drive, and when to keep your hands on the wheel.



