GEOG 30323: Data Analysis & Visualization
2026-09-15
Includes:
Your core Python tools for EDA: NumPy, pandas, and seaborn/matplotlib
import statementstdlib, the standard library that ships with Pythonre for regular expressions; os for operating system functions; random for random-number generation; and many more. Full list: https://docs.python.org/3/library/import numpy as npimport pandas as pdcolleges.csv from the link to your computer/content, Colab’s working directory - so pd.read_csv("colleges.csv") finds itcolleges['median_earn'], or as attributes of the data frame, e.g. colleges.median_earndel statementMake sure you know your column types (dtypes) and levels of measurement before doing analysis!
The mean of a sample (\(\overline{x}\)) is calculated as follows:
\[\overline{x} = \dfrac{x_1 + x_2 + ... + x_n}{n}\]
where \(n\) is the number of elements in the sample.
pandas as data frame methods, e.g. colleges.mean(), colleges.std().describe() will give you back a number of important descriptive stats at once.describe()matplotlibseaborn: extension to matplotlib to make your graphics look nicer! Standard import: import seaborn as sns.pandaspandas and seabornGEOG 30323 | Data Analysis & Visualization