What kind of decision agents does an AI-native org's C-suite need
I built three seemingly unrelated products in a row — Boss, MBA Brand, OAF — and only by the third did I realize they share the same skeleton. This is about that skeleton, and why an AI-native org's C-suite needs it.
Three products, one accidental discovery
In just the past month, I built three AI products in a row, each for a completely different audience:
- Boss (bossagent.cc) — a strategy OS for the CEO. Throw it a strategic question and it convenes a multi-juror deliberation, producing a verdict with attribution checks.
- MBA Brand (mbabrand.com) — a brand-influence audit for the CMO. Give it a brand name and it runs multi-dimensional research, then has a panel of jurors score it independently.
- OAF (oaf.world) — an investment-research workbench for the CFO. Covering the US / A-share / HK markets, it hands the numbers to deterministic tools and the narrative to the large model.
Different domains, different users, different data sources. But by the third one I stopped — because I realized I was writing the same thing for the third time.
A common enemy: gut calls
All three roles make high-stakes judgments every day: do we invest in this direction, is this brand actually strong, is this company worth the price. These judgments have long leaned on the same word — gut feel.
"Gut feel" has two fatal failure modes:
- Single-perspective bias. You ask only one person, even the most senior one, and his blind spot becomes your blind spot.
- Conclusions can't be attributed. A gut call of "looks good," and half a year later it collapses — you can't tell whether the judgment was wrong, the evidence was incomplete, or the world changed.
Three different roles, all afflicted by the same disease. So the cure should be the same set too.
The same skeleton
What I kept rebuilding across all three products was the same pipeline:
Research → Deliberate → Score → Attribute
parallel jurors alone multi-lens 30/90/365
gather ev. unseen to ea. falsifiable timeline review Four steps; let me explain why for each.
A panel of jurors, not a single analyst
A single model (or a single person) delivering a conclusion is still "single perspective" at heart. So every judgment step is decided not by one agent but by a panel of jurors, each holding a stance, scoring independently and required to argue the opposing side — one watches strategy, one watches evidence, one plays devil's advocate on purpose. Before deliberation the jurors can't see each other, avoiding herd effects. What comes out isn't one opinion but a map of disagreement: which conclusions are consensus and which are hotly contested, at a glance.
Scorable and falsifiable, not "I think it's fine"
"This brand is very strong" can't be reviewed later. "This brand scores 8.8 on category naming but only 5.2 on genuine signal, because its online reach concentrates in a single channel" — that can. Every conclusion is required to give a score and a falsification point: under what conditions it would be counted wrong. Judgments can now be rebutted rather than merely believed.
Versioned + attribution-checked, not one-and-done
A judgment is timestamped, frozen into an immutable snapshot, and pre-set with attribution checkpoints at 30 / 90 / 365 days. When they come due, the system returns to compare: did the original prediction hit? Hits / falsifications are written into the scoreboard. Judgment thus becomes, for the first time, an accumulable asset — you're not just making decisions, you're building a backtestable curve of your own judgment.
Anti-fabrication is a hard constraint, not an option
The most dangerous thing about a large model is that it can invent a number with a straight face. So all three products make "anti-fabrication" a machine-level constraint: use only public first-hand / real data; if it can't be obtained, mark it N/A and leave it blank, never patch it up for you. OAF simply hands all numbers to deterministic tools and lets the model write only the narrative; MBA wires juror citations into a CI hard gate, verifying word-for-word whether they're actually in the corpus. Better to say less than to say wrong.
Dual readable, by humans and machines
The same judgment must be readable by humans (HTML report / dashboard) and callable by agents (MCP / API). Because in an AI-native org, the downstream of a judgment includes both people and other agents.
What changes is only the jurors and the data sources
Once you see the skeleton, the differences between the three products are actually quite superficial:
| Who the jurors are | Data sources | |
|---|---|---|
| Boss | dimensional doctrine + anchor conviction | strategic questions, internal docs |
| MBA | persona mental models (Fu Sheng / Jobs…) | brand signals, public sentiment |
| OAF | long-short debate + deterministic tools | three-market quotes and filings |
The skeleton didn't change by a line. That's also why they can add up to one set, rather than three isolated tools.
Why "AI-native org"
Take it one level higher.
An AI-native org is one that hands more and more of its daily operations — writing code, running research, doing customer support, watching data — to agents. When the execution layer already runs on traceable agents, the fact that only top-level judgment still lives in the boardroom's intuition becomes the most glaring gap.
What the C-suite decision agents aim to fill is exactly this layer: to let the CEO / CMO / CFO's core judgments grow on the same "researchable, scorable, reviewable" foundation as the rest of the org. Not to replace the human who calls the shot — the person still presses the button in the end — but to turn the research, deliberation, scoring, and attribution that precede the call into infrastructure that both humans and agents can read and revisit.
The more AI-native an org is, the more indispensable this layer of infrastructure becomes.
Compressing the skeleton into a single mark
Finally, I made a mark for all of this — three arcs, one dot, one seal:
Three graphic elements, each carrying a creed:
- A seal. The site's cinnabar red (
#b14b3a) is the color of seal paste to begin with, and "stamping a seal" is the ultimate act of decision. This suite is about versioned freezing — each verdict generates an immutable snapshot — and a seal naturally carries that meaning: once stamped, it is an immutable version. - Three broken arcs forming a C. The three arcs are three seats (CEO / CMO / CFO); broken, because the jurors score independently, unseen to each other; read together as a C, they are deliberated into one, one skeleton. The opening faces right — an open judgment, not a closed loop of self-justification.
- A single dot of ink at the center. The only element in the whole image that isn't cinnabar. It is the anchor: the person who presses the button in the end. The machine orbits the human rather than replacing the human's call — exactly what the previous section said, drawn into the mark.
The full design derivation, geometric parameters, variants, and usage guidelines are on the Logo and brand page.
Boundaries
Let me be honest. This thing is no oracle:
- The jurors are AI-deduced perspectives, not a real board; they force you to think it through, but they don't take responsibility for you.
- Scores are relative and corpus-dependent; when the corpus changes the score should change — which is precisely why it must be versioned.
- Attribution checks need time; before the 30/90/365 days are up, the scoreboard is still empty.
Even so, moving judgment from "gut calls" to "a traceable pipeline" is itself already an upgrade. The rest, leave to time to backtest.
All three entry points are here: Boss · CEO, MBA Brand · CMO, OAF · CFO, or enter from the C-suite feature page.