Multivariate data exploration

GEOG 30323: Data Analysis & Visualization

Dr. Kyle Walker

2026-09-22

Why visualize data?

The greatest value of a picture is when it forces us to notice what we never expected to see.

  • Tukey (1977) quoted in Yau (2013)

For discussion

https://www.nytimes.com/interactive/2018/08/04/upshot/up-birth-age-gap.html

  • What type of chart is this?
  • What does it suggest?

Exploring data visually

Source: Wilke, Fundamentals of Data Visualization, Ch. 5 (CC BY-NC-ND)

Our schedule:

  • Current activities: data exploration through visualization with common chart types
  • Weeks 11-13: deep dive into data visualization
    • More complex chart types
    • How to customize your seaborn plots
    • Best practices in data visualization
    • Interactive web-based graphics
    • Maps!

Exploratory chart types

  • Comparing categories: bar chart, dot plot
  • Part-to-whole: pie chart
  • Change over time: line chart
  • Connections and relationships: scatter plot

Many, many more in these categories - these are just our focus for today!

Python and the web

  • A brief aside: With Python, data on the web is at your fingertips (our topic for Week 10)
  • This week, you will get a preview
import pandas as pd

mx_csv = "https://raw.githubusercontent.com/walkerke/geog30323/main/data/mexico.csv"
mx = pd.read_csv(mx_csv)
mx.head()
                  name  gdp_pc  gdp_growth   pri   sec   ter  other  pct_65plus
0       Aguascalientes   214.5        -1.9   4.3  30.3  64.7    0.7         7.2
1      Baja California   233.5         0.3   5.2  29.5  61.4    3.9         7.1
2  Baja California Sur   211.8         3.8   5.9  18.8  74.8    0.5         6.6
3             Campeche   497.3        -6.7  17.6  18.8  63.4    0.2         7.9
4              Chiapas    64.4         2.2  30.9  18.0  50.9    0.2         6.0

Comparing categories

How about sorting our data?

mx_sorted = mx.sort_values(by = 'gdp_pc', ascending = False)
mx_sorted.head()
                name  gdp_pc  gdp_growth   pri   sec   ter  other  pct_65plus
3           Campeche   497.3        -6.7  17.6  18.8  63.4    0.2         7.9
6   Ciudad de México   422.5         2.2   0.6  13.5  84.4    1.5        12.0
18        Nuevo León   326.6         3.4   1.3  33.5  64.7    0.5         8.0
7           Coahuila   275.0        -0.5   2.2  37.6  59.4    0.8         7.9
25            Sonora   265.0        -0.3   8.9  28.3  61.0    1.8         8.8

Bar charts

Source: Pew Research Center, “Young Adults and the Future of News” (Dec 2025)

Bar charts

  • Length or height of bars proportional to data values, allowing for comparisons between categories
  • The value axis of bar charts must start at zero!!!
  • Recommendation: sort your data values for ease of interpretation

Bar chart with non-zero origin

Source: @WhiteHouse on X, Jan 30, 2026 (3M views). Community Note: “The vertical axis only shows values in the range 80.2Mt to 81.8Mt. This makes the 1% increase appear much larger than it actually is.”

Bar charts in Python

import seaborn as sns
sns.set(style = "darkgrid")

mx.plot(x = 'name', y = 'gdp_pc', kind = 'bar')

Bar charts in seaborn

sns.barplot(x = 'gdp_pc', y = 'name', data = mx_sorted)

Dot plots

Source: Pew Research Center, “Americans’ Social Media Use 2025” (Nov 2025)

Dot plots

  • Can be preferable to bar charts - values determined by position along axis rather than bar heights
  • In turn, zero origin not strictly necessary (though consider the context)
  • Sorted data also preferable for dot plots

Dot plots in seaborn

sns.stripplot(x = 'gdp_pc', y = 'name', data = mx_sorted)

Part-to-whole

  • Categories in relationship to the entire population of values
  • Examples: pie chart, waffle chart, 100% bar chart, tree map
  • Must sum to 100%!

Pie charts in Python

zac = mx[mx.name == 'Zacatecas']
zac = zac.drop(['name', 'gdp_pc', 'gdp_growth', 'pct_65plus'], axis = 1)
zac = zac.squeeze()
zac.name = 'Zacatecas'
zac.plot(kind = 'pie', figsize = (6, 6))

Problems with pie charts

Source: Times Now (India) broadcast graphic, June 2020, via WTF Visualizations

Problems with pie charts

Source: Data to Viz

Line charts

Source: Pew Research Center, “Teens, Social Media and AI Chatbots 2025” (Dec 2025)

Line charts in seaborn

dfw = pd.read_csv('https://raw.githubusercontent.com/walkerke/geog30323/main/data/pct_college.csv')

sns.lineplot(x = "year", y = "pct_college", 
             hue = "county", data = dfw)

Scatter plots

  • Question: how do the values in two columns covary?
  • Scatter plot: each observation represented by a point; position along x axis dictated by one column value; position along y axis dictated by other column value
  • Regression line: visual representation of estimated statistical relationship between X and Y

Scatter plots

Source: Georgetown CEW, “Major Payoff” (Feb 2026; ACS 2021-23)

Scatter plots in seaborn

sns.scatterplot(x = "pri", y = "gdp_pc", data = mx)

Scatter plots in seaborn

  • Also available in the lmplot and regplot functions
sns.lmplot(data = mx, x = 'pri', y = 'gdp_pc')

Correlation

  • Correlation coefficient: statistical representation of how two samples covary; ranges between -1 (negative correlation) and +1 (positive correlation)
  • In pandas: .corr()
  • Beware of spurious correlations! http://tylervigen.com/spurious-correlations
mx['pri'].corr(mx['gdp_pc'])
-0.4880127198490067