All Things Trust · 2025–present
Two products have come out of All Things Trust: the Trust Stack credibility diagnostic, and Decision Lens, the analysis product built on its methodology.
Case study 01 · 2025
The Trust Stack scoring system
A diagnostic that scores how credible an organisation looks to the people and AI systems evaluating it. One score, five dimensions, evidence behind every number, and a named human who can overrule the model before anything ships.
- My role
- Partner & Principal Architect
- What I owned
- Scoring hierarchy, evaluation pipeline, human review instrument, report product
- Status
- Shipped, running live diagnostics
Situation
Companies invest heavily in how they present themselves, and have almost no way to see how credible that presentation actually is to the consumers weighing it. Trust gets measured by proxy, through sentiment, reviews, and brand tracking, and none of those says whether a claim can be believed.
The tools that did exist measured reputation, or measured AI visibility. Neither answers the question a consumer is actually asking: is the evidence behind this claim reachable at the moment I have to decide? The Trust Stack diagnostic scores a company’s presence the way that consumer would experience it: what can be found, understood, and checked.
My role, precisely
I didn’t create the Trust Stack framework. It was here before I was, five dimensions already named and defined. What I built is everything that turns it into a scored, evidenced, repeatable diagnostic: the four-level scoring hierarchy, the evaluation pipeline, the human review instrument, and the report clients receive.
How it works
Four stages, and a human in the third one
A score is only useful if you can see how it was produced, so the pipeline is disclosed inside the report itself, on the same screen as the number. Three kinds of judgement feed it: heuristic checks, LLM analysis, and human review.
Architecture
Only the bottom level is scored; everything above it is arithmetic
Every level sits on the same 1.0–10.0 scale. Attributes are the only things the model and the reviewer score directly; signals, dimensions, and the overall number are all derived from them. So when a client challenges a 6.8, the answer is never “the model felt that way.” It’s the list of attributes underneath, and who scored each one.
Rubric
Every band resolves to a word, not just a number
A 4.2 means nothing to a marketing director. “Your claims are Referenced but not Substantiated” means something they can act on this week. Each dimension gets its own four-state ladder, so the score always maps to a condition a human can argue with.
| Dimension | Weak 1.0–3.9 | Fair 4.0–5.9 | Good 6.0–7.9 | Excellent 8.0–10.0 |
|---|---|---|---|---|
| Provenance | Unclear | Identified | Attributable | Traceable |
| Verification | Unsubstantiated | Referenced | Substantiated | Corroborated |
| Transparency | Hidden | Disclosed | Clear | Answerable |
| Coherence | Fragmented | Aligned | Integrated | Unified |
| Resonance | Generic | Relevant | Context-aware | Responsive |
The hard calls
Three calls I’d defend, and one open question
01 · Weighting
Provenance and Verification at 25% each; Transparency 20%; Coherence and Resonance 15%. In plain terms: who is behind this and can it be checked matter about 1.7× more than whether it lands well.
What it costs: a well-sourced but tone-deaf experience outscores a warm, resonant one with no evidence behind it. I think that’s correct for a credibility instrument. It is also the first thing a client pushes back on, so the weights are shown on the same screen as the score rather than buried in an appendix.
02 · A scale, not a verdict
Four named bands instead of pass/fail. A scale shows a client where they stand and lets them sequence the fixes; what it gives up is the blunt clarity of a verdict, which executives find easier to act on.
03 · Core vs. amplifier signals
Signals are typed CORE or AMP, so a dimension
isn’t a flat mean of its parts. A missing accountability signal should
hurt more than a thin participatory-engagement signal, and typing the
signals is how that judgement gets written into the maths.
04 · Where the human sits, and the anchoring problem
The review sheet shows the reviewer the model’s score for a signal before they enter their own. That is deliberate: it makes reconciliation fast enough to be economically viable on every engagement rather than a research exercise.
It also anchors the human toward the model. I don’t have a full answer for it yet, and I’d rather say so here than have someone spot it later. The way to measure it is to blind-rate a sample: hide the model score for a subset of signals and compare the distribution of human scores against the shown-score condition. If the drift is small, the throughput gain is free. If it isn’t, the fix is to blind the CORE signals and leave the amplifiers anchored.
How I knew it worked
The evaluation, which is most of the actual work
- Multi-model cross-validation. Applied reviews run across OpenAI, Anthropic and Google models to refine definitions, reduce overlap between dimensions, and harden the scoring logic. Reducing overlap is discriminant validity work: if Coherence and Resonance move together across every subject, one of them isn’t earning its place.
- An empirically located human boundary. The same testing was used to identify where human interpretation remains necessary. Where the model families kept disagreeing with each other is where the human review now sits; I didn’t have to guess.
- Paired scores at signal level. The review sheet records the model’s score and the human’s score side by side, with a free-text note, which turns rater agreement into a number we can track over time.
- Ground truth where ground truth exists. Representation Accuracy and Category Classification Accuracy compare AI-generated descriptions against verified facts and official category definitions, not against a judgement call.
- Construct validity, in progress and labelled as such. Academic and industry partnerships are now testing whether the five dimensions are genuinely distinct, measurable, and consequential for behaviour. Until that lands, the framework is well-grounded but not validated, and the product says so.
Grounding
The scoring model was built against the existing literature rather than intuition: Fogg & Tseng on computer credibility (1999), Sundar’s MAIN model (2008), Metzger, Flanagin & Medders on heuristic credibility evaluation (2010), Lee & See on trust in automation (2004), and Liao & Sundar on designing for responsible trust in AI systems (2022), alongside C2PA Content Credentials and the Content Authenticity Initiative for the provenance side.
What broke
Draft — Andrew to complete
One honest paragraph goes here, and its absence is the difference between a case study and a brochure. Candidates from the artefact itself:
- A dimension score hiding its own spread: in one report Provenance rolled up to 7.3 from signals ranging 4.0 (Traceability) to 9.2 (Accountability). Did a client ever act on the 7.3 and miss the 4.0?
- Where the three model families disagreed most, and what you changed because of it.
- A signal you cut because it couldn’t be scored reliably.
Outcome
A shipped diagnostic: a scored report with an executive summary, dimension breakdown, strengths and vulnerabilities, a cited-source evidence layer, editable inputs, and export.
The part I’m proudest of is less visible. The report documents its own limits inside the product: how it was built, what it scores, and an explicit statement of what it does not score: not reputation, not security, not legal compliance, not product efficacy. Brand recognition and marketing spend are not treated as evidence of credibility. Spelling that out costs us some easy sales conversations, and it keeps the score from being stretched to cover things it was never built to measure.
Live reports are confidential. The diagnostic methodology, scoring engine, signal taxonomy and implementation model are proprietary to All Things Trust; everything on this page is either publicly documented or my own account of decisions I made. Happy to walk through a redacted report on a call.