The architecture in one sentence
A WhatsApp API does not think, it transports. It receives the customer's message, notifies your system and delivers whatever reply you send. The intelligence lives in your backend, which talks to the model provider of your choice. Keeping the two apart is what lets you switch models tomorrow without touching the WhatsApp integration.
- The customer sends a message. The API fires
messages.receivedto your webhook. - The endpoint puts the event on a queue and returns 200 right away.
- A worker loads the history for that number and the relevant business data.
- The worker calls the model with instructions, context and the new message.
- The reply goes back to the customer with
messages.sendText, or the case is handed to an agent.
Step 2 is not a detail. A model can take several seconds to respond, and the webhook has a 30-second timeout. Calling the AI inside the webhook request triggers retries and duplicate replies.
Worker pseudocode in Node
import { DApi } from 'd-api-sdk'
const dapi = new DApi({ apiKey: process.env.DAPI_KEY })
export async function handleMessage({ sessionId, data }) {
if (data.fromMe || data.is_group) return
if (await alreadyProcessed(data.id)) return
const phone = data.from.jid.split('@')[0]
const conversation = await loadConversation(sessionId, phone)
if (conversation.assignedToHuman) {
return forwardToAgentInbox(conversation, data)
}
conversation.append({ role: 'user', content: data.message })
// your LLM provider; D-API is not part of this step
const reply = await llm.generate({
system: BUSINESS_RULES,
context: await fetchCustomerFacts(phone),
messages: conversation.lastTurns(12),
})
if (reply.needsHuman) {
await conversation.assignToHuman(reply.reason)
return dapi.messages.sendText({ sessionId, to: phone, text: 'Let me connect you with someone from our team. One moment.' })
}
conversation.append({ role: 'assistant', content: reply.text })
await dapi.messages.sendText({ sessionId, to: phone, text: reply.text })
}Functions like loadConversation and llm.generate are yours. What comes from D-API is the incoming event and the outgoing send. For more on the client, see the Node.js SDK.
Context and history: what to send the model
The model remembers nothing between calls. Everything it knows about the conversation is what you send. Three layers are usually enough:
- Fixed rules: who the assistant is, what it can and cannot answer, tone of voice, when to bring in a human.
- Customer facts: open orders, current plan, next appointment. Fetched from your database on every message, not left to the model's memory.
- Recent turns: a window of the latest messages. Long conversations get expensive and slow; summarize what came before instead of sending everything.
Store history keyed by connection and number. In a product with many customers, each with their own WhatsApp, the sessionId identifies which customer the conversation belongs to, and mixing the two is a data leak.
Handing off to a human without losing the customer
Every chatbot needs an exit. The most common triggers are an explicit request ("I want to talk to a person"), repeated frustration, sensitive topics like cancellations or complaints, and low confidence from the model itself.
When the case goes to a human, flag the conversation and stop answering with AI until the agent hands it back. Messages the agent sends from the number's phone arrive as messages.received with fromMe: true, which helps detect that someone took over. If your product is already a support platform, this is the same flow described in WhatsApp API for helpdesks.
Latency and cost: where the time goes
| Step | What weighs | How to reduce it |
|---|---|---|
| Fetching context | Database queries and calls to internal systems | Cache the facts that rarely change |
| Calling the model | Prompt size and model size | Short history window, smaller model for triage |
| Media | Audio transcription, image understanding | Process only when the flow needs it |
| Sending the reply | One HTTP call | Little to optimize; the API responds fast |
On cost, the number that matters is per conversation, not per message. The model charges for the volume of text processed. On the WhatsApp side, D-API charges per connection, not per message, so a chatty bot does not raise your channel bill. The details are on the pricing page.
What to check before going to production
- Hallucination: price, deadlines, balance and status always come from your system, never from generated text.
- Prompt injection: the customer can write "ignore your rules". Treat the message as data and limit the actions the AI can trigger.
- GDPR and LGPD: disclose the automation, minimize what goes to the model, set a retention period for the history and know where the provider processes the data.
- Outbound volume: a bot that replies to everything, including another bot, can get stuck in a loop. Cap replies per conversation and ignore groups if groups are not part of the use case.
If you would rather build the flow with visual tools instead of code, compare with Typebot or with automation flows.
