Measuring semantic similarity between two texts — 4 approaches compared

Cosine similarity, BERTScore, LLM-as-judge, or a calibrated alignment score? Here is how each approach performs on real-world AI output evaluation.

Why semantic similarity is harder than it looks

You have two pieces of text. You want to know how well they "match." Sounds simple. It is not.

The naive approach — cosine similarity on bag-of-words vectors — fails the moment synonyms or paraphrasing appear. *"The customer is unhappy"* and *"The buyer expressed dissatisfaction"* should score high. Bag-of-words says they barely overlap.

This article compares four approaches that handle real semantic matching, with tradeoffs.

Approach 1: Cosine similarity on embeddings

**How it works:** Convert each text to a dense vector with an embedding model (OpenAI `text-embedding-3-small`, Cohere, sentence-transformers), then compute cosine similarity.

**Strengths:** Fast, cheap, works offline if you self-host the model.

**Weaknesses:** Treats meaning as a single point. Loses nuance around emotion, intent, and brand voice. Two texts about completely different topics can score 0.7+ if they share vocabulary.

Approach 2: BERTScore

**How it works:** Compares contextual BERT embeddings token-by-token, then aggregates with precision/recall.

**Strengths:** Better than cosine for paraphrase detection. Established benchmark in NLP literature.

**Weaknesses:** Expensive at scale. Still anchored to surface form — a sarcastic match scores like a sincere match.

Approach 3: LLM-as-judge

**How it works:** Send both texts to GPT-4 / Claude / Gemini and ask "how well do these match on a 0-100 scale?"

**Strengths:** Captures intent, tone, and reasoning. Can explain its score.

**Weaknesses:** Slow, expensive, non-deterministic. Same input can score 73 one minute and 81 the next. No structured output without prompt engineering.

Approach 4: Calibrated alignment scoring

**How it works:** A purpose-built scoring API that uses an LLM with structured tool-calling to return a calibrated 0-100 score plus a per-dimension breakdown (semantic, emotional, intent, brand voice).

**Strengths:**

**Example:**


curl -X POST https://resonanceengine.p.rapidapi.com/evaluate \
  -H "X-RapidAPI-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": "I need help fast",
    "target": "Sure, here is a 12-step explainer",
    "mode": "intent"
  }'

Returns:


{
  "score": 23,
  "reasoning": "The target ignores the urgency signal in the source...",
  "mode": "intent"
}

Low score correctly flags the mismatch — the user wanted speed, the assistant delivered length.

Which to pick

| Use case | Best fit |

|----------|----------|

| Search/retrieval at scale | Cosine on embeddings |

| Academic NLP benchmarks | BERTScore |

| One-off qualitative review | LLM-as-judge |

| Production AI output evaluation | Calibrated alignment scoring |

Try it

[ResonanceEngine on RapidAPI](https://rapidapi.com/fredrikssonanton01/api/resonanceengine) — calibrated 0-100 alignment scoring across semantic, emotional, intent, and brand dimensions. One API call, structured response.

Loading interactive experience…