← Real buildsCoolbackCustomer care · case file

PRD_Coolback-Order-Processing-Final

Turns inbound B2B order emails into validated, human-checked, ERP-ready orders — matching customers and delivery addresses against Airtable and filing every outcome back into the shared mailbox as a category tag. — 5,388 tasks (5,371 completed, 9 failed) in the 30 days to 2026-08-10 — the highest-volume agent on the platform; 2,721 tasks in the 14 days to 2026-07-13

In production79 nodes in the graph8 steps drawn
—
no measured bill on either rail
1
Classify email intent (order vs not) and tag the Outlook category
2
Validate the customer type; extract order-head details
3
Fetch customer records from Airtable; match the delivery address; tag exceptions ('PLZ nicht gefunden', 'Adresse nicht gefunden')
4
Attach the customer ID; extract the order line items
5
Merge everything and pause for the human-in-the-loop check
6
Generate the order CSV and upload via SFTP for ERP import
7
Archive the order PDF to DocuWare; append the log row to Google Sheets
8
Tag the mailbox 'Bestellung verarbeitet'

Where the money goes, per run

Step typeCountRateSubtotal
Code / integration steps46$0.025$1.15
Model calls · standard tier12$0.1$1.2
Model calls · mid tier20$0.2$4
Document parses1$0.2$0.2
Rate-card estimate, standard tier$6.55

Full graph at rate card: 46 code/integration x $0.025 = $1.15; 12 std-model x $0.10 = $1.20; 20 mid-model x $0.20 = $4.00; 1 intake parse x $0.20 = $0.20; total $6.55 — a ceiling only, since the 11 exclusive rule-based conditions mean one lane runs per task. Edge-walk of the costliest single lane (order-processed path): 24 nodes = 13 code ($0.325) + 4 std ($0.40) + 6 mid ($1.20) + parse ($0.20) = $2.13; a 'Keine Bestellung' early exit is roughly $0.45. The code/integration 46 = 35 tool nodes (14 Outlook tag, 6 Sheets append, 3 each DocuWare/CreateCsv/Airtable/S3/SFTP) + 11 rule-based conditions; std 12 = 11 Flash prompt nodes + 1 GPT4_1_MINI; mid 20 = 14 GPT5 + 6 GEMINI_3_1_PRO (GPT-5 = mid tier is EVIDENCED by the platform rate table in 03-projects/27-cx/01-planning/overview.md, which prices '判断·中档(GPT-5 · Gemini 2.5 Pro)' at 2 credits). No measured per-run figure exists anywhere in the repo (catalog cost_basis 'none'); the honest fallback per the catalog meta-note is the menu price plus the at-cost band. Dropping the 6 mid nodes on the happy lane to std would cut it from $2.13 to ~$1.53.

Models seen in the graph: preferredModel: GEMINI_3_FLASH x46, GPT5 x14, GEMINI_3_1_PRO x6, GPT4_1_MINI x1; all 11 conditionNodes rule_based with llmModel GPT4_1_MINI; fallback chains include GEMINI_3_PRO, GPT4_1, CLAUDE_4_SONNET, GPT5_2, GPT5_4, BEDROCK_CLAUDE_SONNET_4_5

Graph read: 04-agents/coolback/coolback-order-processing-final/agent.json (79 nodes, graph a8276c55; agent.new.json is the same graph id, also 79 nodes)

What a copycat build should know

  • Mailbox as state machine: 14 MicrosoftOutlookAction_AddMessageCategory terminals write German status tags ('Verarbeitung', 'Bestellung verarbeitet', 'PLZ nicht gefunden', 'Adresse nicht gefunden', 'Keine Bestellung') back to the shared Outlook box — the mailbox IS the case tracker (documented as the state-machine-tags topology pattern, 03-projects/beam-agent-design/01-topology-patterns/05-state-machine-tags.md).
  • Convergence is hand-built: the extract->match->format->CSV->SFTP->DocuWare->Sheets writer lane is cloned ~3x (3 CreateCsv, 3 SftpUploader, 3 DocuWare, 6 SheetAppend) because Beam has no merge node — every branch terminates into its own duplicated writers (03-projects/beam-agent-design/06-roadmap-watch-list.md names Coolback's 7-way convergence explicitly).
  • All 11 conditions are rule_based (deterministic routing); judgment is concentrated in extractors (OrderHead/OrderProduct), matchers (address -> customer ID against Airtable), and dedicated HITL machinery — 'DataMerger/DataPreparationforHITL' GPT nodes prepare the human check and 6 'HiTLFeedback' GPT nodes consume the reviewer's verdict after the pause.
  • Tier assignment is deliberate, not uniform: email-intent and order-head extraction sit on GEMINI_3_1_PRO, spreadsheet formatting and HITL feedback on GPT5, address/product matching on Flash — the expensive models guard the fields that reach the ERP CSV.

Lessons from the client record

  • Scope lesson (02-design-decisions/accuracy-target-vs-autonomy-level.md, quoted in beam-agent-design/GUIDE.md): Coolback shipped 120 days unbilled in P5 doing custom HITL development that should have been P3 scope, because the DoD said 'agent runs autonomously' instead of 'agent meets X% accuracy on Y dataset' — a new build must define accuracy-on-dataset as the DoD.
  • Graph flags are proven unusable for production status: this agent shows isPublished/isEverExecuted all-false against 5,388 probed completions in 30 days — it is the calibration case behind the catalog meta-note's rule to trust only task-probe numbers.
  • Highest-volume agent on the platform (5,388 tasks, 9 failed, 30d to 2026-08-10) yet the catalog carries no per-run cost — high volume makes the mid-tier node choices the dominant cost lever.
  • When the platform Merge node ships, the duplicated-writer branches should be refactored (06-roadmap-watch-list.md 'Merge Node for Branch Convergence', status not started).

What it is made of

Case creationMicrosoft OutlookAirtableSFTP (ERP import)DocuWareGoogle SheetsS3

04-agents/coolback/coolback-order-processing-final/agent.json · 04-agents/coolback/coolback-order-processing-final/README.md · 04-agents/_probe/task_ranking_2026-08-10.json · 03-projects/beam-agent-design/GUIDE.md · 03-projects/beam-agent-design/02-design-decisions/accuracy-target-vs-autonomy-level.md · 03-projects/beam-agent-design/06-roadmap-watch-list.md · 03-projects/beam-agent-design/01-topology-patterns/05-state-machine-tags.md · 03-projects/27-cx/01-planning/overview.md · 03-projects/alloyed-cx/04-outputs/production-cx-table-zh.md · /Users/zhichaoli/Documents/GitHub/solution-intelligence/content/agent-catalog.json