All Things Trust · 2025–present

Two products have come out of All Things Trust: the Trust Stack credibility diagnostic, and Decision Lens, the analysis product built on its methodology.

Case study 01 · 2025

The Trust Stack scoring system

A diagnostic that scores how credible an organisation looks to the people and AI systems evaluating it. One score, five dimensions, evidence behind every number, and a named human who can overrule the model before anything ships.

My role
Partner & Principal Architect
What I owned
Scoring hierarchy, evaluation pipeline, human review instrument, report product
Status
Shipped, running live diagnostics

Situation

Companies invest heavily in how they present themselves, and have almost no way to see how credible that presentation actually is to the consumers weighing it. Trust gets measured by proxy, through sentiment, reviews, and brand tracking, and none of those says whether a claim can be believed.

The tools that did exist measured reputation, or measured AI visibility. Neither answers the question a consumer is actually asking: is the evidence behind this claim reachable at the moment I have to decide? The Trust Stack diagnostic scores a company’s presence the way that consumer would experience it: what can be found, understood, and checked.

My role, precisely

I didn’t create the Trust Stack framework. It was here before I was, five dimensions already named and defined. What I built is everything that turns it into a scored, evidenced, repeatable diagnostic: the four-level scoring hierarchy, the evaluation pipeline, the human review instrument, and the report clients receive.

How it works

Four stages, and a human in the third one

A score is only useful if you can see how it was produced, so the pipeline is disclosed inside the report itself, on the same screen as the number. Three kinds of judgement feed it: heuristic checks, LLM analysis, and human review.

The four-stage Trust Stack evaluation pipeline Stage one, automated scan applying heuristic checks to brand-controlled pages. Stage two, LLM-assisted external evidence review linked to specific claims. Stage three, signal-level human reconciliation of the model's scores. Stage four, weighted synthesis into dimension scores and one overall score. 01 Automated scan Scoped crawl and heuristic checks of included pages. Brand-controlled media 02 Evidence review LLM-assisted. External proof linked to specific claims. Cited sources 03 Human reconciliation Signal-by-signal review of the model’s scores. Reviewer of record 04 Synthesis Weighted roll-up, findings, recommendations. One score
Stage 03 is the one most scoring systems skip, and the only stage where a named person can disagree with the output before it rolls up.

Architecture

Only the bottom level is scored; everything above it is arithmetic

Every level sits on the same 1.0–10.0 scale. Attributes are the only things the model and the reviewer score directly; signals, dimensions, and the overall number are all derived from them. So when a client challenges a 6.8, the answer is never “the model felt that way.” It’s the list of attributes underneath, and who scored each one.

The four-level Trust Stack scoring hierarchy Over one hundred observable attributes are scored, resolving into more than twenty credibility signals, which resolve into five dimension scores, which are weighted into one overall score. Dimension weights are Provenance 25 percent, Verification 25 percent, Transparency 20 percent, Coherence 15 percent, Resonance 15 percent. 100+ Observable attributes Scored directly by reviewer and model. 20+ Credibility signals Typed CORE or AMP; they don’t weigh alike. 5 Dimensions Provenance, Resonance, Coherence, Transparency… 1 Overall score Weighted average of the five dimensions. Dimension weights Provenance  25% Verification  25% Transparency  20% Coherence  15% Resonance  15% Bar widths are proportional to weight.
Provenance and Verification carry half the score between them. That’s a deliberate position on what credibility means, not something that fell out of tuning.

Rubric

Every band resolves to a word, not just a number

A 4.2 means nothing to a marketing director. “Your claims are Referenced but not Substantiated” means something they can act on this week. Each dimension gets its own four-state ladder, so the score always maps to a condition a human can argue with.

Trust Stack rubric states by dimension and band
Dimension Weak 1.0–3.9 Fair 4.0–5.9 Good 6.0–7.9 Excellent 8.0–10.0
ProvenanceUnclearIdentifiedAttributableTraceable
VerificationUnsubstantiatedReferencedSubstantiatedCorroborated
TransparencyHiddenDisclosedClearAnswerable
CoherenceFragmentedAlignedIntegratedUnified
ResonanceGenericRelevantContext-awareResponsive

The hard calls

Three calls I’d defend, and one open question

01 · Weighting

Provenance and Verification at 25% each; Transparency 20%; Coherence and Resonance 15%. In plain terms: who is behind this and can it be checked matter about 1.7× more than whether it lands well.

What it costs: a well-sourced but tone-deaf experience outscores a warm, resonant one with no evidence behind it. I think that’s correct for a credibility instrument. It is also the first thing a client pushes back on, so the weights are shown on the same screen as the score rather than buried in an appendix.

02 · A scale, not a verdict

Four named bands instead of pass/fail. A scale shows a client where they stand and lets them sequence the fixes; what it gives up is the blunt clarity of a verdict, which executives find easier to act on.

03 · Core vs. amplifier signals

Signals are typed CORE or AMP, so a dimension isn’t a flat mean of its parts. A missing accountability signal should hurt more than a thin participatory-engagement signal, and typing the signals is how that judgement gets written into the maths.

04 · Where the human sits, and the anchoring problem

The review sheet shows the reviewer the model’s score for a signal before they enter their own. That is deliberate: it makes reconciliation fast enough to be economically viable on every engagement rather than a research exercise.

It also anchors the human toward the model. I don’t have a full answer for it yet, and I’d rather say so here than have someone spot it later. The way to measure it is to blind-rate a sample: hide the model score for a subset of signals and compare the distribution of human scores against the shown-score condition. If the drift is small, the throughput gain is free. If it isn’t, the fix is to blind the CORE signals and leave the amplifiers anchored.

How I knew it worked

The evaluation, which is most of the actual work

  • Multi-model cross-validation. Applied reviews run across OpenAI, Anthropic and Google models to refine definitions, reduce overlap between dimensions, and harden the scoring logic. Reducing overlap is discriminant validity work: if Coherence and Resonance move together across every subject, one of them isn’t earning its place.
  • An empirically located human boundary. The same testing was used to identify where human interpretation remains necessary. Where the model families kept disagreeing with each other is where the human review now sits; I didn’t have to guess.
  • Paired scores at signal level. The review sheet records the model’s score and the human’s score side by side, with a free-text note, which turns rater agreement into a number we can track over time.
  • Ground truth where ground truth exists. Representation Accuracy and Category Classification Accuracy compare AI-generated descriptions against verified facts and official category definitions, not against a judgement call.
  • Construct validity, in progress and labelled as such. Academic and industry partnerships are now testing whether the five dimensions are genuinely distinct, measurable, and consequential for behaviour. Until that lands, the framework is well-grounded but not validated, and the product says so.

Grounding

The scoring model was built against the existing literature rather than intuition: Fogg & Tseng on computer credibility (1999), Sundar’s MAIN model (2008), Metzger, Flanagin & Medders on heuristic credibility evaluation (2010), Lee & See on trust in automation (2004), and Liao & Sundar on designing for responsible trust in AI systems (2022), alongside C2PA Content Credentials and the Content Authenticity Initiative for the provenance side.

What broke

Draft — Andrew to complete

One honest paragraph goes here, and its absence is the difference between a case study and a brochure. Candidates from the artefact itself:

  • A dimension score hiding its own spread: in one report Provenance rolled up to 7.3 from signals ranging 4.0 (Traceability) to 9.2 (Accountability). Did a client ever act on the 7.3 and miss the 4.0?
  • Where the three model families disagreed most, and what you changed because of it.
  • A signal you cut because it couldn’t be scored reliably.

Outcome

A shipped diagnostic: a scored report with an executive summary, dimension breakdown, strengths and vulnerabilities, a cited-source evidence layer, editable inputs, and export.

The part I’m proudest of is less visible. The report documents its own limits inside the product: how it was built, what it scores, and an explicit statement of what it does not score: not reputation, not security, not legal compliance, not product efficacy. Brand recognition and marketing spend are not treated as evidence of credibility. Spelling that out costs us some easy sales conversations, and it keeps the score from being stretched to cover things it was never built to measure.

Live reports are confidential. The diagnostic methodology, scoring engine, signal taxonomy and implementation model are proprietary to All Things Trust; everything on this page is either publicly documented or my own account of decisions I made. Happy to walk through a redacted report on a call.