← Real buildsCMI (via Alloyed; anonymized as 'Adenva' in the sandbox workspace)Media · case file

SBX_Adenva Social Pacing

Backfills campaign pacing trackers from ad-platform exports: code computes all the spend and impression figures, AI touches only new ad-set names, and the write-back changes exactly three manual cells per row — never a formula cell. — 18 sandbox tasks (14 completed, 4 failed) in the 30 days to 2026-08-10 — test runs, not client volume; the media line docs state none of the five media scenarios has daily production volume yet

Built7 nodes in the graph6 steps drawn
—
no measured bill on either rail
1
Parse platform export CSVs and tracker state; flag stale dates before anything runs (code)
2
Deterministically derive YTD, yesterday, and last-7-days spend and impressions per ad set (code)
3
Match ad-set names against the persistent mapping table; only new names go to AI (code)
4
AI maps new or drifted names to template rows with confidence and evidence
5
Code writes only the 3 manual cells per row by row type, skipping actualized and formula rows
6
Send an audit-style run note and completion email reporting every exception with its reason

Where the money goes, per run

Step typeCountRateSubtotal
Code / integration steps4$0.025$0.1
Model calls · top tier2$0.4$0.8
Other1—$0
Rate-card estimate, standard tier$0.9

As configured (both prompt nodes preferredModel CLAUDE_4_SONNET; Claude Sonnet = top tier on the 2026-02-13 rate card per ama-solution/03-projects/25-media/01-planning/overview.md and 14-agent-modularization/01-planning/two-axes.md): 1 entry/trigger x $0.025 + 4 Code Executor x $0.025 = $0.125, + 2 top-tier judgments x $0.40 = $0.80, + $0 parse (CSV parsing runs inside a Code Executor, dodging the $0.20 parse floor) = $0.925/run = 9.25 credits. At the menu's standard-tier quoting convention the same shape is $0.125 + 2 x $0.10 = $0.325/run. No measured figure exists to compare (catalog cost_basis 'none'; the probe shows 18 sandbox tasks, no credit-ledger entry). The catalog's five mapped menu modules + media intake sum to $0.555/run (menu.json: intake $0.13 + collect-files $0.05 + normalise-format $0.05 + resolve-names $0.15 + check-pacing $0.025 + write-back-and-draft $0.15) — the built graph undercuts the menu at standard tier because parse and extraction are folded into code, but shipping it with Sonnet pinned would bill $0.925, 1.7x the menu price; downgrading the two judgment nodes to a standard-tier model is what closes that gap.

Models seen in the graph: CLAUDE_4_SONNET preferred on all 6 execution nodes, fallbacks GEMINI_3_PRO, GPT4_1 (seen in toolConfiguration of every node). Rate-card canon classes Claude Sonnet as top tier ($0.40/judgment), so as-built the 2 judgment nodes bill top, not the standard tier the menu quotes.

Graph read: ama-solution/04-agents/alloyed/sbx-adenva-social-pacing/agent.json — parsed the full 157KB graph JSON (python, whole file); 7 nodes, chain Entry → Parse CSVs → Derive windows → Match vs mapping table → AI Map new names → Write-set → Run note/email; cross-checked node-by-node against ama-solution/03-projects/25-media/02-resources/research/cmi-pacing-graph.md (same 7-node reading).

What a copycat build should know

  • Mapping-table-first name resolution: code indexes the persistent mapping table and resolves known export names deterministically; only unseen names reach the LLM, and proposals with confidence >= 0.8 are both written and emitted as mapping_updates so the table grows and LLM usage shrinks toward zero over runs. The table travels in the input payload each run — no Beam memory tool (isMemoryTool false everywhere).
  • Spend never reaches a model: CSV parsing, gzip handling, and all arithmetic (date-serial y*372+m*31+d, YTD/yesterday/last-7 per ad set) are Code Executor; the LLM sees only name strings and row labels. Parsing in code also avoids the $0.20/cycle parse floor.
  • Guarded write-set: exactly 3 manual cells per row at manual_cells addresses; row_type branch (media rows get spend, fee rows impressions-only — fee spend stays formula-derived); skip/actualized rows, missing targets, confidence < 0.8, and unmatched names each produce a typed exception with reason; a blocking export-date-mismatch flag refuses stale exports before anything runs; +/-10% budget-vs-campaign pacing flag.
  • Exit is an audit: the run-note/email prompt composes from summary_facts only with 'never invent, drop, or soften an exception'; note that no node has eval criteria (isEvaluationEnabled false on all 7), so accuracy rests entirely on the code guards.

Lessons from the client record

  • Media has no production volume on any of its five scenarios — this graph logged 18 sandbox tasks (14 completed, 4 failed) in the 30-day probe; run-rate proof is borrowed from the Americana daily suite and proves workflow shape, not media experience (25-media overview + catalog meta note).
  • The real client runs 17 pacing workbooks and write-back is the heaviest module of the five media cases: 3 cells per row, formula cells never touched (25-media/01-planning/overview.md five-case table).
  • Graph flags isPublished/isEverExecuted are proven unusable (coolback: all-false flags vs 5,388 probed completions); production status comes only from the task probe (overview 2026-08-28 correction).
  • The old '$0.12/node' pricing is deprecated: an integration node and a Sonnet judgment differ 16x, so model tier moves the bill more than node count — the menu quotes standard tier and top tier multiplies judgments by 4 (overview pricing section).

Reusable fragments cut from this agent

  • mediaParse platform export CSVs and tracker state; flag stale dates before anything runs03-projects/25-media/02-resources/module-library/01-收集文件.md#1
  • mediaMatch ad sets against the persistent mapping table; only new names go to the AI mapper03-projects/25-media/02-resources/module-library/03-判断名称.md#1
  • mediaMap only new/drifted ad-set names to template rows with confidence and evidence; never guess03-projects/25-media/02-resources/module-library/03-判断名称.md#2
  • mediaDeterministically derive YTD, yesterday and last-7-days spend/impressions per ad set03-projects/25-media/02-resources/module-library/04-核对投放进度.md#1
  • mediaWrite-set for the workbook: 3 manual cells per row by row type, skip actualized/hidden rows, 10 percent pacing checks, exceptions with reasons03-projects/25-media/02-resources/module-library/07-写回与起草.md#1
  • mediaAudit-log style run note plus completion email; reports every exception with its reason03-projects/25-media/02-resources/module-library/07-写回与起草.md#2

What it is made of

Collect filesNormalise formatResolve namesCheck pacingWrite back and draftAd-platform CSV exportsExcel pacing trackers (17 workbooks at the real client)email

ama-solution/04-agents/alloyed/sbx-adenva-social-pacing/agent.json · ama-solution/03-projects/25-media/02-resources/research/cmi-pacing-graph.md · ama-solution/03-projects/25-media/01-planning/overview.md · ama-solution/03-projects/14-agent-modularization/01-planning/two-axes.md · ama-solution/03-projects/25-media/02-resources/module-library/ (fragment 出处 lines) · solution-intelligence/content/menu.json · solution-intelligence/content/agent-catalog.json