Model card
Compass model card
Version 0.1.0. Generated 2026-10-02 by compass export. Every number below comes from outputs/.
Intended use
Compass suggests which students might benefit from a check-in, and says why in plain language. A school chooses how many check-ins its staff can handle; Compass returns that many students, ranked by need, each with one to three reasons and a suggested support. The output is always "Check-in suggested" or "No check-in suggested". It never labels a student.
Compass is a prototype. It has been tested only on synthetic data and on one small public dataset. It has not been validated on Virginia students. See USE_POLICY.md for prohibited uses.
Inputs
Attendance, behavior, and course indicators only. No race, ethnicity, sex, income or lunch status, disability, English learner status, address, or family background, ever. The code rejects any file that contains those columns.
| Input | Weight (per unit, log-odds) | Typical student (median) | Direction |
|---|---|---|---|
Days absent this grading period (absent_rate) | 11.028 | 0.075 | more means more need |
Suspensions or serious incidents (suspensions) | 0.189 | 0.0 | more means more need |
Core courses failing (core_failures) | 0.000 | 0.0 | more means more need |
GPA this grading period (gpa) | -0.782 | 3.01 | more means less need |
Credits behind on-time pace (credits_behind) | 0.026 | 0.0 | more means more need |
GPA change since last grading period (gpa_change) | -0.091 | 0.0 | more means less need |
| Intercept | -0.623 |
Weights fixed at zero by the sign constraints: Core courses failing. In this synthetic data the other course indicators (GPA and credits behind) already carry that signal, and the constraint stops the model from giving it a backwards weight.
Model
Logistic regression with sign constraints: every weight is fixed in advance to point the sensible way (more absences can never lower a score). Fitted with bounded L-BFGS and a small L2 penalty. Each published weight is the whole model.
Check-in line used by the website demo: a score of 0.272 or higher, which is where the top 15% of synthetic test students begin.
Synthetic evaluation
20,000 synthetic ninth graders in 30 schools (seed 2026). Built-in structural inequality: groups are unevenly spread across lower-resource schools, and recorded suspensions are inflated for Group B (x1.3) and Group C (x1.7) for the same behavior. Test set: 6,000 students. Capacity: 900 check-ins (15%).
| Model | AUC | Precision at capacity | False alarm rate | Miss rate | Largest false alarm gap |
|---|---|---|---|---|---|
| Compass | 0.785 | 46.2% | 9.7% | 57.7% | 6.1 points |
| DEWS-style baseline (uses group and income) | 0.775 | 43.0% | 10.2% | 60.7% | 8.0 points |
| School-level only | 0.661 | 29.6% | 12.6% | 73.0% | 23.8 points |
False alarm rate by group (share of students who graduated on time but were flagged):
| Model | Group A | Group B | Group C |
|---|---|---|---|
| Compass | 7.7% | 10.1% | 13.7% |
| DEWS-style baseline (uses group and income) | 7.3% | 11.6% | 15.3% |
| School-level only | 5.2% | 13.4% | 29.0% |
Audit result: AUDIT FAILED: Compass's false alarm rate ranges from 7.7% (Group A) to 13.7% (Group C), a gap of 6.1 percentage points (limit 5). Compass does not meet its own limit on this synthetic data, so it is not ready to ship. The gap comes mostly from unequal school resources, which shape attendance and grades, not from group membership, which Compass never sees.
Proxy check: with and without suspensions
Average recorded suspensions vs. true incidents per student: Group A 0.22 vs 0.23; Group B 0.34 vs 0.263; Group C 0.512 vs 0.298.
| Compass | AUC | Precision at capacity | Largest false alarm gap |
|---|---|---|---|
| With suspensions | 0.785 | 46.2% | 6.1 points |
| Without suspensions | 0.783 | 46.2% | 6.1 points |
The suspensions weight is small (0.189 per suspension), so dropping it barely moves the gap here. The lesson still holds: removing demographic columns did not remove the gap, because other data carries the inequality.
Public data check: UCI Student Performance
UCI Student Performance (Cortez, CC BY 4.0): 649 students at two schools in Portugal. A sanity check on real data, not proof. Only sex and address are available as group attributes. Outcome: final grade g3 below 10 (a failing grade). Test set: 195 students. Capacity: 29.
By sex:
| Model | AUC | Precision at capacity | Largest false alarm gap |
|---|---|---|---|
| Compass | 0.957 | 75.9% | 0.7 points |
| DEWS-style baseline (uses sex, address, family background) | 0.939 | 69.0% | 3.9 points |
| School-level only | 0.682 | 31.0% | 7.0 points |
Audit passed: Compass's false alarm rate ranges from 4.0% (Female) to 4.7% (Male), a gap of 0.7 percentage points (limit 5).
By address:
| Model | AUC | Precision at capacity | Largest false alarm gap |
|---|---|---|---|
| Compass | 0.957 | 75.9% | 3.0 points |
| DEWS-style baseline (uses sex, address, family background) | 0.939 | 69.0% | 7.2 points |
| School-level only | 0.682 | 31.0% | 18.8 points |
Audit passed: Compass's false alarm rate ranges from 3.4% (Urban) to 6.4% (Rural), a gap of 3.0 percentage points (limit 5).
With 195 test students, these results are noisy. Treat them as a sanity check.
Limits
- Synthetic data reflects our assumptions, documented in
compass/synth.py. Real schools will differ. - The UCI data is small, from Portugal, and only has sex and address as group attributes.
- Several fairness definitions cannot all hold at once when groups have different underlying rates (Kleinberg, Mullainathan, and Raghavan 2016; Chouldechova 2017). Compass prioritizes equal false alarm rates and reports the rest.
- A list does not create counselor time. The school view exists to show where resources are short.
- Built by students. Needs review by professional data scientists, educators, and legal advisors before any real use.