Learn how string and vector distances answer different similarity questions, and why raw-count Euclidean distance tracks document length.
A library volunteer receives two boxes of index cards. One box has typed names with small spelling errors, and the other has paragraphs from speeches.
The word “similar” is doing two jobs. Matching a mistyped name is not the same task as comparing the vocabulary of two paragraphs.
A distance measure turns a comparison into a number. Small distance means close. The hard part is choosing a distance that matches the question.
TipWhat you will learn
This lesson shows how to:
compare edit distance and Jaro-Winkler distance for names;
build paragraph vectors from word counts;
compare cosine, Euclidean, and Jaccard distance;
test whether a distance mostly tracks text length; and
explain why the same data can give different rankings.
Compare strings directly
String distance compares characters. Levenshtein distance counts edits: insertions, deletions, and substitutions. Jaro-Winkler distance gives extra credit to strings that share an early prefix, which often helps with names.
Two string distances rank the same candidate names differently
Recorded
Candidate
Levenshtein
Jaro-Winkler
Levenshtein rank
Jaro-Winkler rank
Jonson
Johnson
1
0.0381
1
2
Jonson
Jansen
2
0.2000
4
6
Martha
Marta
1
0.0333
2
1
Martha
Marsha
1
0.0778
3
5
Stevenson
Stephenson
2
0.0726
5
4
Stevenson
Steven
3
0.0667
6
3
For Stevenson, Levenshtein picks Stephenson because it needs fewer edits. Jaro-Winkler picks Steven because the shared opening is strong. A name-matching project has to choose the error pattern it cares about.
Build paragraph vectors
A vector is a row of numbers. In a document-feature matrix, each row is a paragraph and each column is a word feature. The entry says how many times that word appears in that paragraph.
Paragraph vectors built from the inaugural-address corpus
Item
Value
Paragraphs
1377
Speeches
60
People
40
Features
9327
Paragraph pairs
947376
The matrix has 1,377 paragraph rows and 9,327 word columns. Each paragraph becomes a point in a high-dimensional space.
Compare three vector distances
Euclidean distance measures straight-line distance between count vectors. Cosine distance compares direction, so multiplying a paragraph vector by the same constant leaves its cosine direction unchanged. Jaccard distance uses word presence and ignores repeated counts.
Raw-count distance compared with paragraph length difference
Distance
Correlation with length difference
Euclidean distance
0.951
Cosine distance
-0.268
Jaccard distance
0.275
kable( same_speech_examples,col.names =c("First paragraph", "Second paragraph", "Speech", "First length", "Second length", "Length difference", "Euclidean", "Cosine", "Jaccard"),caption ="Long and short paragraphs from the same speech can be far apart under Euclidean distance",row.names =FALSE)
Long and short paragraphs from the same speech can be far apart under Euclidean distance
First paragraph
Second paragraph
Speech
First length
Second length
Length difference
Euclidean
Cosine
Jaccard
1841-Harrison-p10
1841-Harrison-p22
1841-Harrison
984
69
915
147.46
0.330
0.936
1841-Harrison-p10
1841-Harrison-p25
1841-Harrison
984
80
904
145.87
0.304
0.942
1841-Harrison-p10
1841-Harrison-p11
1841-Harrison
984
94
890
142.01
0.229
0.924
The Euclidean correlation with length difference is 0.951, a strong warning that raw-count distance is mostly measuring paragraph size. The cosine correlation is negative, -0.268, because longer paragraphs share many ordinary words with other long paragraphs. That matters for this corpus: nineteenth-century paragraphs are longer than later paragraphs, so a length-shaped method can echo the era pattern seen in the corpus lessons.
Same-speech paragraph cosine distance compared with shuffled speech labels
Observed
Null mean
Null 5th pct
Null 95th pct
p-value
Replicates
Alternative
0.569
0.577
0.573
0.581
0.003
1000
less
A distance formula will fill a matrix for any paragraphs it receives. The permutation check asks whether paragraphs from the same speech are closer than a label shuffle would produce. Here they are, but the raw-count Euclidean result is still dominated by length.
Figure 1: Length difference and distance for a deterministic sample of paragraph pairs.
The plot turns the correlation into a visible warning. A long paragraph and a short paragraph can look far apart under raw-count Euclidean distance even when they come from the same speech.
What to remember
String distance compares character sequences.
Vector distance compares numeric representations of text.
Levenshtein and Jaro-Winkler can rank the same name candidates differently.
On these paragraph counts, Euclidean distance correlates 0.951 with length difference.
Cosine distance compares direction, which reduces the length problem for raw counts.
The volunteer keeps two labels on the boxes: spelling repair for names, representation choice for paragraphs.