The dataset
The Farrlandia Census
Farrlandia, counted in full: 10,000 residents with their demographics, exposures, biomarkers, and health outcomes. We built the census from a documented causal structure, so you can run any analysis and then check your answer against the process that produced the data. The interactive explorer works from a 1,000-person random sample of this census. Download the full file and analyze it in whatever software you have.
farr <- read.csv("https://epidemiologymatters.com/farrlandia/census/farrlandia-census.csv")
table(farr$smoking, farr$lung_cancer)
prop.table(table(farr$cvd))
import delimited "https://epidemiologymatters.com/farrlandia/census/farrlandia-census.csv", clear tab smoking lung_cancer, row cs cvd heavy_alcohol
import pandas as pd
farr = pd.read_csv("https://epidemiologymatters.com/farrlandia/census/farrlandia-census.csv")
pd.crosstab(farr.smoking, farr.lung_cancer, normalize="index")
GET DATA /TYPE=TXT /FILE="farrlandia-census.csv" /DELIMITERS="," /FIRSTCASE=2. * Download the file first, then point FILE= at it. CROSSTABS /TABLES=smoking BY lung_cancer /CELLS=ROW.
Codebook
The 32 Variables
Binary variables are coded 1 = yes, 0 = no. There are no missing values: this is a census, and every resident answered every question.
| Variable | Type | Description |
|---|---|---|
| id | integer | Resident identifier (10001–99998). |
| age | years | Age at census, 18–79. |
| sex | female / male | Sex. |
| district | text | Residential district: Snow (industrial), Nightingale, Hill, Doll, Rose (most affluent). |
| ses | 1–5 | Socioeconomic position, 1 = lowest, 5 = highest. |
| education | text | Highest education: primary, secondary, tertiary. |
| occupation | text | Coal miner (Hill Colliery), factory worker (Snow), farmhand, service worker, office worker, student, retired, unemployed. |
| water_source | text | Household water: “snow pump” (the old pump serving District Snow’s poorest streets) or “municipal”. |
| laid_off_factory | 0/1 | Lost their job when the Vulcan Works in Snow closed, two years before the census. |
| insured | 0/1 | Has health insurance. |
| smoking | 0/1 | Current smoker. |
| pack_years | continuous | Cumulative smoking dose; 0 for non-smokers. |
| air_pollution | 0/1 | High residential air-pollution exposure (concentrated in Snow and Hill). |
| poor_diet | 0/1 | Diet quality below recommended. |
| physical_inactivity | 0/1 | Below activity guidelines. |
| heavy_alcohol | 0/1 | Heavy episodic drinking (more common under 40). |
| social_isolation | 0/1 | Socially isolated. |
| bmi | kg/m² | Body mass index. |
| systolic_bp | mmHg | Systolic blood pressure. |
| family_history_cvd | 0/1 | First-degree family history of cardiovascular disease. |
| cvd | 0/1 | Prevalent cardiovascular disease (≈7%). |
| depression | 0/1 | Prevalent depression (≈9%). |
| depression_2yr_ago | 0/1 | Depression two years before the census, before the Vulcan Works closed. |
| lung_cancer | 0/1 | Lung cancer, ever diagnosed (≈3%). |
| type2_diabetes | 0/1 | Prevalent type 2 diabetes (≈8%). |
| type2_diabetes_diagnosed | 0/1 | Diabetes that a clinician has actually diagnosed. Compare with the row above; the two differ by district. |
| resp_infection_past_year | 0/1 | Respiratory tract infection, past year (≈18% overall; higher in Hill; see teaching point E). |
| injury_past_year | 0/1 | Injury requiring medical attention, past year (≈11%). |
| clinic_visit_past_year | 0/1 | Attended the Farrlandia clinic in the past year (≈36%). See teaching point D before conditioning on this variable. |
| followup_years | years | Person-years of follow-up in the Farrlandia cohort, 4.0–10.0. |
| household_id | text | Shared dwelling (H0001–H4207). Household members live in one district, draw the same water, and tend toward similar socioeconomic position. Useful as a clustering level in multilevel models. |
| workplace_id | text | Employer group (W0001–) for working residents: colliery crews of about 15, factory shifts of about 25, farm teams, offices. Empty for retired, student, and unemployed residents. |
For Simulation
The Contact Network
Who meets whom in Farrlandia. Every resident's contacts come as an edge list in three layers: household (everyone in a shared dwelling), workplace (colliery crews, factory shifts, farm teams, offices), and community (neighbourhood acquaintances). Together with the census it supports transmission models, agent-based simulation, and network analysis on a population whose causal structure is fully documented.
Download the contact network (CSV)
The layers hold 10,056 household ties, 41,048 workplace ties, and 39,786 community ties, a mean of 18 contacts per resident. The network is deterministic: it is generated with the census from the same fixed seed, so an analysis written today runs unchanged next year. To see these files at work without writing code, the Trial Grounds runs an in silico trial and an epidemic on them in your browser.
R · igraph
edges <- read.csv("https://epidemiologymatters.com/farrlandia/census/farrlandia-contacts.csv")
library(igraph)
g <- graph_from_data_frame(edges, directed = FALSE)
summary(g)
Python · networkx
import pandas as pd, networkx as nx
edges = pd.read_csv("https://epidemiologymatters.com/farrlandia/census/farrlandia-contacts.csv")
G = nx.from_pandas_edgelist(edges, "person_a", "person_b", edge_attr="layer")
print(G)
For instructors
Teaching Points
We generated the census from a known causal structure. Each of the findings below is in the data and will appear when students run the analysis.
A · Confounding
Smoking is associated with cardiovascular disease, but smoking is also patterned by age and socioeconomic position, and both shape CVD on their own. Have students compare crude and adjusted estimates and explain what changed.
B · A masked effect
Crudely, heavy drinkers have the same CVD risk as everyone else. Stratify by age and the harm appears. The reason: heavy drinking is concentrated in the young, whose baseline risk is low. Confounding hid the harm here.
C · Interaction
The effect of smoking on lung cancer is roughly twice as large among residents with high air-pollution exposure as among those without. Two component causes are completing the same sufficient cause, which is Chapter 11 working in the data.
D · Selection (Berkson’s bias)
Restrict any analysis to clinic attendees and strange things happen: type 2 diabetes appears to protect against depression, an association that does not exist in the full census. Clinic attendance is a collider: both conditions bring people through its doors.
E · Structure and place
Respiratory infection runs between 15% and 19% in four districts, and 28% in Hill, home of the Hill Colliery. Miners carry the heaviest burden. Geography and occupation pattern disease here, not chance: ask students what a “district effect” is actually made of.
F · A true null
District Snow’s poorest households draw water from the old Snow Pump. Crudely, pump users have more CVD, but the pump is causally inert: its users are simply poorer. Adjust for socioeconomic position, district, and age, and the association goes away. The crude estimate reflects poverty, not the water.
G · A natural experiment
The Vulcan Works in Snow closed two years before the census. Depression among the workers it laid off rose from 12% to 24%; among the workers it kept, nothing moved. The depression_2yr_ago column lets students build the two-by-two themselves and estimate the effect of job loss as a difference in differences.
H · Ascertainment
The Nightingale Clinic opened eighteen months before the census. Diagnosed diabetes in Nightingale now runs half again higher than in any other district, while true prevalence is flat. The clinic created no disease; it found what the dispensaries were missing. Ask students which of the two numbers a health ministry would see.
Instructor’s answer key: expected values
Values your students should approximately reproduce (risk ratios; Mantel-Haenszel for adjusted):
| Point | Analysis | Expected |
|---|---|---|
| A | Smoking → CVD, crude / age-adjusted | ≈2.4 / ≈2.3 |
| B | Heavy alcohol → CVD, crude / age-adjusted | ≈0.9 / ≈1.5 |
| C | Smoking → lung cancer, RR by pollution stratum | ≈3 (low) vs ≈9 (high) |
| D | Diabetes → depression, full census / clinic only | ≈0.9 / ≈0.5 |
| E | Respiratory infection: Hill vs other districts; miners vs non-miners | ≈28% vs 15–19%; RR ≈2.7 |
| F | Snow Pump → CVD, crude / SES-district-age adjusted | ≈1.5 / ≈1.0 |
| G | Depression, laid off vs kept on, before and after the closure | DiD ≈+13 pp |
| H | Diagnosed diabetes, Nightingale vs elsewhere (true prevalence flat) | ≈6% vs ≈4% |
The census is simulated (fixed seed, generator documented in the site repository), so these values are stable across downloads. Sampling variability applies only to subsamples your students draw themselves.
Provenance
Where This Data Comes From
Farrlandia is simulated. We generated every resident from a documented data-generating process with a fixed random seed, so no real person appears in this file and any answer the data gives can be checked against the code that produced it. Use the census freely in courses, workshops, and problem sets, with attribution to Epidemiology Matters (Keyes, Abba-Aji & Galea, Oxford University Press).