Skip to content
Compass

Model card

Compass model card

Version 0.1.0. Generated 2026-10-02 by compass export. Every number below comes from outputs/.

Intended use

Compass suggests which students might benefit from a check-in, and says why in plain language. A school chooses how many check-ins its staff can handle; Compass returns that many students, ranked by need, each with one to three reasons and a suggested support. The output is always "Check-in suggested" or "No check-in suggested". It never labels a student.

Compass is a prototype. It has been tested only on synthetic data and on one small public dataset. It has not been validated on Virginia students. See USE_POLICY.md for prohibited uses.

Inputs

Attendance, behavior, and course indicators only. No race, ethnicity, sex, income or lunch status, disability, English learner status, address, or family background, ever. The code rejects any file that contains those columns.

InputWeight (per unit, log-odds)Typical student (median)Direction
Days absent this grading period (absent_rate)11.0280.075more means more need
Suspensions or serious incidents (suspensions)0.1890.0more means more need
Core courses failing (core_failures)0.0000.0more means more need
GPA this grading period (gpa)-0.7823.01more means less need
Credits behind on-time pace (credits_behind)0.0260.0more means more need
GPA change since last grading period (gpa_change)-0.0910.0more means less need
Intercept-0.623

Weights fixed at zero by the sign constraints: Core courses failing. In this synthetic data the other course indicators (GPA and credits behind) already carry that signal, and the constraint stops the model from giving it a backwards weight.

Model

Logistic regression with sign constraints: every weight is fixed in advance to point the sensible way (more absences can never lower a score). Fitted with bounded L-BFGS and a small L2 penalty. Each published weight is the whole model.

Check-in line used by the website demo: a score of 0.272 or higher, which is where the top 15% of synthetic test students begin.

Synthetic evaluation

20,000 synthetic ninth graders in 30 schools (seed 2026). Built-in structural inequality: groups are unevenly spread across lower-resource schools, and recorded suspensions are inflated for Group B (x1.3) and Group C (x1.7) for the same behavior. Test set: 6,000 students. Capacity: 900 check-ins (15%).

ModelAUCPrecision at capacityFalse alarm rateMiss rateLargest false alarm gap
Compass0.78546.2%9.7%57.7%6.1 points
DEWS-style baseline (uses group and income)0.77543.0%10.2%60.7%8.0 points
School-level only0.66129.6%12.6%73.0%23.8 points

False alarm rate by group (share of students who graduated on time but were flagged):

ModelGroup AGroup BGroup C
Compass7.7%10.1%13.7%
DEWS-style baseline (uses group and income)7.3%11.6%15.3%
School-level only5.2%13.4%29.0%

Audit result: AUDIT FAILED: Compass's false alarm rate ranges from 7.7% (Group A) to 13.7% (Group C), a gap of 6.1 percentage points (limit 5). Compass does not meet its own limit on this synthetic data, so it is not ready to ship. The gap comes mostly from unequal school resources, which shape attendance and grades, not from group membership, which Compass never sees.

Proxy check: with and without suspensions

Average recorded suspensions vs. true incidents per student: Group A 0.22 vs 0.23; Group B 0.34 vs 0.263; Group C 0.512 vs 0.298.

CompassAUCPrecision at capacityLargest false alarm gap
With suspensions0.78546.2%6.1 points
Without suspensions0.78346.2%6.1 points

The suspensions weight is small (0.189 per suspension), so dropping it barely moves the gap here. The lesson still holds: removing demographic columns did not remove the gap, because other data carries the inequality.

Public data check: UCI Student Performance

UCI Student Performance (Cortez, CC BY 4.0): 649 students at two schools in Portugal. A sanity check on real data, not proof. Only sex and address are available as group attributes. Outcome: final grade g3 below 10 (a failing grade). Test set: 195 students. Capacity: 29.

By sex:

ModelAUCPrecision at capacityLargest false alarm gap
Compass0.95775.9%0.7 points
DEWS-style baseline (uses sex, address, family background)0.93969.0%3.9 points
School-level only0.68231.0%7.0 points

Audit passed: Compass's false alarm rate ranges from 4.0% (Female) to 4.7% (Male), a gap of 0.7 percentage points (limit 5).

By address:

ModelAUCPrecision at capacityLargest false alarm gap
Compass0.95775.9%3.0 points
DEWS-style baseline (uses sex, address, family background)0.93969.0%7.2 points
School-level only0.68231.0%18.8 points

Audit passed: Compass's false alarm rate ranges from 3.4% (Urban) to 6.4% (Rural), a gap of 3.0 percentage points (limit 5).

With 195 test students, these results are noisy. Treat them as a sanity check.

Limits

  • Synthetic data reflects our assumptions, documented in compass/synth.py. Real schools will differ.
  • The UCI data is small, from Portugal, and only has sex and address as group attributes.
  • Several fairness definitions cannot all hold at once when groups have different underlying rates (Kleinberg, Mullainathan, and Raghavan 2016; Chouldechova 2017). Compass prioritizes equal false alarm rates and reports the rest.
  • A list does not create counselor time. The school view exists to show where resources are short.
  • Built by students. Needs review by professional data scientists, educators, and legal advisors before any real use.