Measuring brand voice consistency with AI: 4 axes that matter
Cosine similarity gives you one opaque number. Real brand voice QA needs four — semantic, emotional, intent, and tone. Here is how to score AI output against your brand bible.
Measuring brand voice consistency with AI: 4 axes that matter
Your AI writes a customer email. Is it "on brand"? Most teams answer with vibes — or worse, with a single cosine similarity score that conflates four very different things.
This post breaks down the four axes that actually matter for brand voice QA, and shows how to score them with one API call.
The problem with one number
Cosine similarity on embeddings gives you a single 0–1 score for "how similar are these two texts?" That is useful for retrieval. It is **terrible** for QA.
Why? Because two texts can be:
- Semantically identical but tonally wrong (formal vs casual)
- Tonally aligned but missing intent (friendly but not actually helpful)
- On-message but off-brand (right idea, wrong voice)
One score cannot tell you which dimension failed.
The four axes
**1. Semantic alignment** — Do both texts say the same thing? (factual content match)
**2. Emotional tone** — Do they convey the same feeling? (warm, urgent, calm, excited)
**3. Intent alignment** — Do they aim for the same outcome? (inform, persuade, reassure, sell)
**4. Brand voice consistency** — Do they sound like the same author? (vocabulary, rhythm, register)
A healthy QA score requires all four to land above your threshold.
Scoring with ResonanceEngine
One endpoint, four scores:
const res = await fetch("https://resonance-engine.p.rapidapi.com/score", {
method: "POST",
headers: {
"X-RapidAPI-Key": process.env.RAPIDAPI_KEY,
"Content-Type": "application/json"
},
body: JSON.stringify({
reference: brandBibleSnippet,
candidate: aiGeneratedEmail
})
});
const { semantic, emotional, intent, brand_voice, total } = await res.json();
// { semantic: 87, emotional: 72, intent: 91, brand_voice: 64, total: 78 }
In this example the AI nailed intent (91) but missed brand voice (64) — probably too formal. Now you know exactly what to fix in the prompt.
Practical use cases
**Pre-send guardrail** — Block any AI-generated message scoring under 70 on brand voice.
**Prompt engineering loop** — Iterate prompts until all four axes hit your threshold consistently.
**Long-term agent memory** — Store user context snippets, retrieve by resonance score for the most contextually relevant memory at inference time. Cheaper and more interpretable than vector DBs for many use cases.
**A/B testing AI versions** — Compare two model outputs against the same reference to see which matches your brand better.
Why not embeddings + custom logic?
You can build this. You will need:
- Embedding pipeline (OpenAI ada or similar — pay per token)
- Sentiment classifier
- Intent classifier (LLM-as-judge, slow + expensive)
- Style classifier (custom training)
ResonanceEngine is one stateless REST call, sub-second, no data retention, $19/mo for 25k requests.
Get started
See the [ResonanceEngine page](/resonance-engine) for live examples and pricing. Free tier on RapidAPI lets you score 100 pairs to evaluate fit.
— Stop shipping off-brand AI. Score everything.
Loading interactive experience…