Julian Teusch

UrbanAI 2026 · ACM SIGSPATIAL workshop · Accepted / In Press

Freeze, Validate, Report.

Auditing Urban Station Plans with Common Evidence

Julian Teusch · Oliver Keszöcze

Compare the plans, not different versions of the evidence. A shared evaluation contract makes station-planning decisions traceable from inputs to the final test report.

A two-point example from the paper

Changing what you score can reverse the ranking.

Plan S₁

500 m
0 m1,000 m

Plan S₂

450 m
400 m500 m

The declared mean ranks S₂ first: 450 m versus 500 m. Lower is better.

Illustrative distances, not city results. Retaining only each plan’s best-served point changes the target: the reported minimum is no longer the declared two-point mean.

The evaluation contract

Same evidence. Explicit priorities. One test report.

01

Freeze

Fix target trips, proxy groups, reporting cells, candidate pools and metric priorities before comparing plans.

02

Validate

Score all seven generators on common validation evidence. Select by worst-group p90 first, then the declared tie-break metrics.

03

Report

Record the selected plan before loading test data. Evaluate that plan once, without choosing again on the test results.

Porto and Chicago

A selection rule, not a universal winner.

The predeclared rule selects IFkCO in Porto and Grid in Chicago. The plots show validation trade-offs; the table reports only the selected plans on held-out test data.

Porto validation: IFkCO is selected; overall and worst-group p90 distances differ across seven methods.
Porto-SES · Validation evidence. Stars mark selected plans; color indicates the cell gap.

Porto-SES

IFkCO

Selected plan

805 mOverall p90
872 mWorst-group p90

Selected plans · test endpoint distances (origin + destination access)

Selected plans · test endpoint distances (origin + destination access)
SettingSelected planOverall p90Worst-group p90
Porto-SESIFkCO805 m872 m
Chicago-HardshipGrid1,404 m1,656 m

Chicago’s primary test uses a later time window on the same day. Metric priorities are choices of the authority; selection by point estimates does not establish statistically reliable superiority.

Post hoc sensitivity check

Candidate sites are part of the decision.

200 stations per plan · changing the candidate pool
Setting300 candidates · primary600 candidates · exploratory
Porto-SESIFkCOGrid
Chicago-HardshipGridPriority

Expanding the candidate set changes the selected generator in both cities. A separate Chicago check that holds the hour fixed across dates selects Priority. These are exploratory checks, not retrospective replacements for the primary analysis.

Scope: auditability within a declared contract, not causal fairness, independently replicated results or stable performance after deployment. Proxy groups, time windows and practical indifference thresholds need explicit justification.

Citation

Julian Teusch and Oliver Keszöcze. Freeze, Validate, Report: Auditing Urban Station Plans with Common Evidence. Accepted at the 4th ACM SIGSPATIAL International Workshop on Advances in Urban-AI (UrbanAI 2026). arXiv:2609.39064.