Computational Analysis of Social Complexity
Fall 2026, Spencer Lyon
Prerequisites
Introduction to Graphs
Strong and Weak Ties
Outcomes
Understand the concept of homophily
Practice working through “by hand” examples of diagnosing homophily
Be prepared to computationally diagnose homophily in a large network
References
Easley and Kleinberg chapter 4 (especially section 4.1)
Datasets
Florentine family relationships: https://
www .cs171 .org /2018 /assets /instructions /lab8 /Lab8 .html
Introduction¶
Main Idea¶
Consider your friends. Do they tend to
Enjoy the same movies, music, hobbies as you?
Hold similar religious or political beliefs?
Come from similar schools, workplaces, or socio-economic settings?
What about a random sample of people in the world?
If you are like me, your answers likely indicate that you have more in common with your friends than you would expect to have with a random sample of people
This concept -- that we are similar to our friends -- is called homophily
Homophily in Graphs¶
In the context of graphs or networks, homophily means that nodes that are connected are more similar than nodes at a further distance in the graph
But what do we mean by more similar?
Idea: We might have common friends.
This is an intrinsic force that led to node formation (e.g. triadic closure)
Alternative: We may share characteristics or properties that are not represented in the graph -- external forces.
Examples: same race, gender, school, employer, sports team, etc.
These external forces are what homophily captures
Context¶
To identify if homophily is active in a network, we must have access to context on top of list of nodes and edges
One way to represent this context would be with a DataFrame in addition to a graph:
One row per node
One column indicating the node identifier (or just use row number)
One column for additional characteristic
Suppose these six people are friends. We record whether each person plays soccer or chess.
Question: We haven’t changed any friendships. Could the characteristic we choose change what we find?
In Figure 1, match each node to its entry in the table. Both characteristics occur in three people, but their cross-type friendships differ.
Figure 1:The friendships stay fixed. Changing the characteristic changes which edges connect different types.
using DataFrames, Graphs, GraphPlot
df1 = DataFrame(
family=[
"Acciaiuoli", "Albizzi", "Barbadori", "Bischeri", "Castellani",
"Ginori", "Guadagni", "Lamberteschi", "Medici", "Pazzi",
"Peruzzi", "Ridolfi", "Salviati", "Strozzi", "Tornabuoni"
],
wealth=[10, 36, 55, 44, 20, 32, 8, 42, 103, 48, 49, 27, 10, 146, 48],
priorates=[53, 65, missing, 12, 22, missing, 21, 0, 53, missing, 42, 38, 35, 74, missing],
)It will be easier to do our homohpily calculations with binary data,
we’ll create new columns,
high_wealthandhigh_powerif the wealth and priorates columns, respectively, are above the column medians
using Statistics
df1[!, :high_wealth] = df1.wealth .> median(df1.wealth)
df1[!, :high_power] = df1.priorates .> median(df1.priorates[.!(ismissing.(df1.priorates))])
df1[ismissing.(df1.priorates), :high_power] .= false
df1
marriages = [
0 0 0 0 0 0 0 0 1 0 0 0 0 0 0
0 0 0 0 0 1 1 0 1 0 0 0 0 0 0
0 0 0 0 1 0 0 0 1 0 0 0 0 0 0
0 0 0 0 0 0 1 0 0 0 1 0 0 1 0
0 0 1 0 0 0 0 0 0 0 1 0 0 1 0
0 1 0 0 0 0 0 0 0 0 0 0 0 0 0
0 1 0 1 0 0 0 1 0 0 0 0 0 0 1
0 0 0 0 0 0 1 0 0 0 0 0 0 0 0
1 1 1 0 0 0 0 0 0 0 0 1 1 0 1
0 0 0 0 0 0 0 0 0 0 0 0 1 0 0
0 0 0 1 1 0 0 0 0 0 0 0 0 1 0
0 0 0 0 0 0 0 0 1 0 0 0 0 1 1
0 0 0 0 0 0 0 0 1 1 0 0 0 0 0
0 0 0 1 1 0 0 0 0 0 1 1 0 0 0
0 0 0 0 0 0 1 0 1 0 0 1 0 0 0
]
g1 = Graph(marriages){15, 20} undirected simple Int64 graphgplot(g1, nodelabel=df1.family)Measuring Homophily¶
Our discussion of homophily so far has been conceptual... let’s make it precise
We’ll compare observed edge frequencies with a random-formation benchmark
This resembles a null-hypothesis comparison from statistics, but we won’t carry out a formal significance test here
Random Homophily¶
Start with a thought experiment: edge formation does not depend on characteristic
We have nodes, and of them have characteristic
The share with this characteristic is
For a simple benchmark, imagine drawing each endpoint’s type independently from these population shares
Both have :
Neither has :
One has and the other does not:
This is our independent-endpoint benchmark. For a finite graph whose edges join distinct nodes, it is an approximation.
Question: Why do we add two terms for a cross-characteristic pair?
Figure 2 shows the four outcomes. What happens to the cross-type probability when one group gets smaller?
Figure 2:The two highlighted outcomes describe opposite orders of drawing the endpoint types. Together they give ; they do not mean we count an undirected edge twice.
Counting Frequencies¶
Now an empirical value...
Let there be edges
Let...
| variable | meaning |
|---|---|
| # edges between 2 | |
| # edges between 2 not | |
| # edges between 1 and 1 not |
Then
We’ll use these 4 numbers to count frequencies of edges between types and non- types
Testing for Homophily¶
We are now ready to compare our network with the benchmark
Recall the two quantities:
: the cross-type probability under the independent-endpoint benchmark
: the observed proportion of cross-characteristic edges
The direction of the difference helps us describe the pattern:
| Comparison | Pattern |
|---|---|
| well above | consistent with inverse homophily |
| near | near the random-formation benchmark |
| well below | consistent with homophily |
Intuition: fewer cross-type edges than the benchmark means more same-type edges
A difference alone does not tell us whether it is statistically significant or whether the characteristic caused the edges to form
Example: high school relationships¶
Recall the graph of romantic relationships between high school students in Figure 3.
Each node is a student; an edge records a romantic relationship during the study’s 18-month observation period.
Question: does this graph exhibit homophily in gender? Why?

Figure 3:Romantic relationships in a high school. Original image reproduced from Easley and Kleinberg, Networks, Crowds, and Markets (2010), Figure 2.7. Study: Bearman, Moody, and Stovel (2004), Chains of Affection: The Structure of Adolescent Romantic and Sexual Networks. Image credited to Mark Newman on the book’s website.
Example: Florentines¶
Let’s work through an example of measuring homophily using the Florentine data
I’ll repeat the data below
df1node_color = map(x -> x ? "blue" : "red", df1.high_wealth)
gplot(g1, nodelabel=df1.family, nodefillc=node_color)Step 1: Counting frequencies¶
First we need to count frequencies for all our characteristics
We’ll do that here
using DataStructuresfunction count_frequencies(vals)
counts = DataStructures.counter(vals)
total = length(vals)
Dict(c => v / total for (c, v) in pairs(counts))
endcount_frequencies (generic function with 1 method)count_frequencies(df1.high_wealth)Dict(
n => count_frequencies(df1[!, n])
for n in names(df1)[4:end]
)Dict{String, Dict{Bool, Float64}} with 2 entries:
"high_wealth" => Dict(0=>0.533333, 1=>0.466667)
"high_power" => Dict(0=>0.666667, 1=>0.333333)Step 2: Counting Edges¶
Next we need to count the number of edges of each type
This step is a bit trickier: we’ll need both the graph and the DataFrame
To not spoil the fun, we’ll leave this code as an exercise on the homework
For now we’ll count by hand
Question: Before looking at the tally, can you find every edge joining a high-wealth family to a low-wealth family?
Use the H/L labels and dashed cross-type edges in Figure 4. Count each marriage once.
Figure 4:The wealth groups use the same threshold as our Julia code: high means strictly above the median of 42. Cross-type marriages account for 12 of the 20 edges.
Let’s compare the marriage network with the benchmark for
high_wealthThere are 7 high-wealth families among the 15 families in our graph
Counting edge types gives:
High–high edges: 5
Low–low edges: 3
Cross edges: 12
Total: edges
The observed cross-edge proportion is
The high-wealth share is
E = ne(g1)
Exy = 12 # cross edges
n_high = 7
N = nv(g1)
px = n_high / N
# compare with the benchmark
2 * px * (1-px), Exy/EOur independent-endpoint benchmark is
The observed cross-edge share is 0.60, about 0.10 higher
This points in the direction of inverse homophily: more cross-wealth marriages than the benchmark predicts
We haven’t established statistical significance, and the comparison alone doesn’t tell us why these families married
Exercise¶
Repeat the counting exercise, but for the high_power characteristic
What do you find? Do you see homophily in this characteristic?