Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
54 commits
Select commit Hold shift + click to select a range
ea4cdd2
Fixed Code for 3 files
kartik-s21 Apr 7, 2026
91d22a7
NameError Resolved
kartik-s21 Apr 8, 2026
c3084c4
Merge branch 'master' into code_fix_unenergy
kartik-s21 Apr 10, 2026
297374c
Merge branch 'master' into code_fix_unenergy
kartik-s21 Apr 12, 2026
34bb357
Merge pull request #1 from kartik-s21/code_fix_unenergy
kartik-s21 Apr 13, 2026
ee1e0b2
Merge branch 'datacommonsorg:master' into master
kartik-s21 Apr 30, 2026
8707bba
Merge branch 'datacommonsorg:master' into master
kartik-s21 May 5, 2026
5527dcd
Merge branch 'datacommonsorg:master' into master
kartik-s21 May 5, 2026
e41a8a7
Merge branch 'datacommonsorg:master' into master
kartik-s21 May 5, 2026
7f17aa3
Resolve merge conflict in energy process script
kartik-s21 May 17, 2026
59bcad2
Merge branch 'datacommonsorg:master' into master
kartik-s21 May 17, 2026
4fc439a
Merge branch 'datacommonsorg:master' into master
kartik-s21 May 28, 2026
9e5417c
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 2, 2026
cb72442
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 4, 2026
d9bb0a6
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 5, 2026
2515f96
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 7, 2026
c099907
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 10, 2026
1359560
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 12, 2026
0dfeecc
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 15, 2026
b09267f
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 17, 2026
776feff
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 18, 2026
4b06212
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jun 30, 2026
8bb164a
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 8, 2026
7ebbdce
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 13, 2026
7eb092e
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 20, 2026
2f9a79b
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 20, 2026
fb658bd
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 22, 2026
29db8fa
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 24, 2026
3026f30
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 29, 2026
7a815af
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 30, 2026
cb02292
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 31, 2026
305fb69
Merge branch 'datacommonsorg:master' into master
kartik-s21 Jul 31, 2026
a146ef9
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 3, 2026
c7ad25e
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 3, 2026
2b81be4
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 4, 2026
5a4e570
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 6, 2026
5b6dbc1
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 17, 2026
ac6de11
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 19, 2026
dd720d8
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 19, 2026
f44bb19
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 24, 2026
fe7698d
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 24, 2026
4472a5e
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 25, 2026
9aaa16b
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 26, 2026
7c56c23
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 27, 2026
49da787
Merge branch 'datacommonsorg:master' into master
kartik-s21 Aug 28, 2026
2e1e382
Automate California School Performance import
kartik-s21 Sep 3, 2026
33f793f
Add validation_config.json, counters, node_mcf, and golden_observations
kartik-s21 Sep 3, 2026
2d7fd94
Address review comments: download error handling, CLI parsing, entity…
kartik-s21 Sep 3, 2026
8a341f7
Fix import name in manifest.json
kartik-s21 Sep 3, 2026
88e22f4
Updated pvmapping and download.py
kartik-s21 Sep 3, 2026
a3772bb
Resolve statvar merge conflicts, use canonical GCS MCF, update golden…
kartik-s21 Sep 3, 2026
10169f3
Merge branch 'master' into ca-school-performance-automation
balit-raibot Sep 7, 2026
cbc5479
Remove .mcf, .gitignore, and run script; relocate unit tests and upda…
kartik-s21 Sep 8, 2026
9576630
Standardize test suite to *_test.py, format with yapf, and update docs
kartik-s21 Sep 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 113 additions & 0 deletions statvar_imports/us_education/california_school_performance/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
# California School Performance (CAASPP) Data Commons Import

## Overview
This directory contains the automated Data Commons import pipeline for California public school academic performance data from the **California Assessment of Student Performance and Progress (CAASPP)** Smarter Balanced Summative Assessments.

The data covers student achievement in **English Language Arts/Literacy (ELA)** and **Mathematics** across California public schools, school districts, counties, and statewide aggregates from 2015 to the present.

- **Primary Source**: California Department of Education (CDE) / Educational Testing Service (ETS)
- **Research Portal**: [https://caaspp-elpac.ets.org/caaspp/ResearchFileListSB](https://caaspp-elpac.ets.org/caaspp/ResearchFileListSB)
- **Import Tier**: Automated StatVar Import

---

## Directory Structure

```
california_school_performance/
├── config/
│ ├── california_school_performance_metadata.csv # Processor configurations (delimiters, headers, output columns)
│ └── california_school_performance_pvmap.csv # Property-Value mappings for subjects, grades, demographics, and metrics
├── counters/
│ └── california_school_performance_counters.csv # Transformation counter summary
├── golden_data/
│ ├── golden_observations.csv # Critical golden entity/place observations
│ └── golden_summary_report.csv # Golden summary report for statvar validation
├── test_data/ # Test fixtures for E2E validation
│ ├── sample_input.txt # Sample raw CAASPP input records
│ ├── sample_state_output.csv # Expected sample CSV observations
│ └── sample_state_output.tmcf # Expected sample TMCF mapping
├── california_school_performance_test.py # Unit tests for download and stream normalization
├── download.py # Automated data fetch and extraction script
├── manifest.json # Data Commons import specification
├── README.md # Documentation and usage guide
└── validation_config.json # Production import validation rules
```


---

## Statistical Variables (StatVars)

The pipeline maps raw test records into **5,966 distinct Statistical Variables** conforming to the Data Commons education schema:

- **Population Type**: `Student`
- **School Subjects**:
- `EnglishLanguageArts` (Test ID: 1)
- `Mathematics` (Test ID: 2)
- **Grade Levels**:
- Grades 3, 4, 5, 6, 7, 8, 11 (`dcid:SchoolGrade3` – `dcid:SchoolGrade11`)
- Grade 13 represents "All Grades combined", mapped to unconstrained student population StatVars without grade constraints (e.g. `Count_Student_EnglishLanguageArts`)
- **Demographic Subgroups (55 groups)**:
- Gender (`Male`, `Female`)
- Race / Ethnicity (`White`, `BlackOrAfricanAmericanAlone`, `Asian`, `Filipino`, `HispanicOrLatino`, `AmericanIndianOrAlaskaNative`, `NativeHawaiianOrOtherPacificIslanderAlone`, `TwoOrMoreRaces`)
- Socioeconomic Status (`EconomicallyDisadvantaged`, `NotEconomicallyDisadvantaged`)
- Intersection of Race × Socioeconomic Status (16 groups)
- Disability Status (`WithDisability`, `NoDisability`)
- English Learner & Fluency Status (`InitialFluentProficient`, `ReclassifiedFluentProficient`, `OnlyEnglish`, `CurrentLearner`, `EverLearned`, `Adult`, `ToBeDetermined`, duration `< 12 months`, duration `>= 12 months`)
- Parent Education Level (`LessThanHighSchoolGraduate`, `HighSchoolGraduateIncludesEquivalency`, `SomeCollegeNoDegree`, `CollegeGraduate`, `GraduateSchoolOrPostGraduate`, `CA_DeclinedToState`)
- Special Student Populations (`Homeless`, `HavingHome`, `Foster`, `NotFoster`, `Migrant`, `NotMigrant`, `FamilyOfArmedForces`, `NotFamilyOfArmedForces`)
- **Metrics**:
- `Total Students Tested with Scores`: Count of students (`measuredProperty: count`)
- `Mean Scale Score`: Mean assessment score (`measuredProperty: assessmentScore`, `statType: meanValue`)
- `Percentage Standard Exceeded`: Educational achievement level `CA_StandardExceeded` (`scalingFactor: 100`, `unit: Percent`)
- `Percentage Standard Met`: Educational achievement level `CA_StandardMet` (`scalingFactor: 100`, `unit: Percent`)
- `Percentage Standard Met and Above`: Educational achievement level `CA_StandardMetAndAbove` (`scalingFactor: 100`, `unit: Percent`)
- `Percentage Standard Nearly Met`: Educational achievement level `CA_StandardNearlyMet` (`scalingFactor: 100`, `unit: Percent`)
- `Percentage Standard Not Met`: Educational achievement level `CA_StandardNotMet` (`scalingFactor: 100`, `unit: Percent`)

---

## How to Run

### 1. Download and Process the Entire Dataset (All Available Years: 2015–2025)
To download and process all years present at source (skipping 2020 when CAASPP was cancelled statewide due to COVID-19):
```bash
# Step 1: Download and normalize data
python3 download.py --years=all

# Step 2: Generate TMCF and CSV observations using stat_var_processor
python3 ../../../tools/statvar_importer/stat_var_processor.py \
--input_data=input_files/sb_ca_all_years_normalized.txt \
--pv_map=config/california_school_performance_pvmap.csv \
--config_file=config/california_school_performance_metadata.csv \
--existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf \
--output_path=output_files/california_school_performance_all_years_output \
--output_counters=counters/california_school_performance_counters.csv
```
This automatically fetches each year, normalizes differences across formats, and runs `stat_var_processor.py` over the consolidated multi-year dataset (`sb_ca_all_years_normalized.txt`), generating **64,000+ observations** across all 10 years.

### 2. Download and Process Specific Years
To fetch and process specific years:
```bash
# Specific range
python3 download.py --years=2015-2024

# Specific single year
python3 download.py --years=2024
```

### 3. Quick Test Run
To download a lightweight test sample and verify the pipeline:
```bash
python3 download.py --test_mode
```

---

## Output Files
The pipeline produces standard Data Commons import artifacts in `output_files/`:
- `california_school_performance_all_years_output.csv`: Complete multi-year observations table (~64,000 rows across 2015–2025)
- `california_school_performance_all_years_output.tmcf`: Template MCF linking columns to Data Commons schema nodes
- `california_school_performance_{YEAR}_output.csv`: Per-year individual observation tables

Original file line number Diff line number Diff line change
@@ -0,0 +1,119 @@
# Copyright 2025 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Unit tests for California School Performance (CAASPP) download and normalization."""

import io
import os
import sys
import unittest
from unittest.mock import patch

# Ensure import directory is in sys.path
_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
if _SCRIPT_DIR not in sys.path:
sys.path.insert(0, _SCRIPT_DIR)

import download


class CaliforniaSchoolPerformanceDownloadTest(unittest.TestCase):

def test_discover_year_urls_known_years(self):
"""Verify known year URL mapping retrieval."""
urls_2015 = download.discover_year_urls(2015)
self.assertIsNotNone(urls_2015)
self.assertEqual(urls_2015,
('sb_ca2015_all_csv_v3.zip', 'sb_ca2015_1_csv_v3.zip'))

urls_2024 = download.discover_year_urls(2024)
self.assertIsNotNone(urls_2024)
self.assertEqual(urls_2024,
('sb_ca2024_all_csv_v1.zip', 'sb_ca2024_1_csv_v1.zip'))

@patch.object(download, 'probe_url_exists', return_value=False)
def test_download_parse_years(self, _):
"""Verify year string parsing logic for single, list, range, and 'all'."""
self.assertEqual(download.parse_years('2023'), [2023])
self.assertEqual(download.parse_years('2023, 2024'), [2023, 2024])
self.assertEqual(download.parse_years('2021-2023'), [2021, 2022, 2023])
all_years = download.parse_years('all')
self.assertIn(2015, all_years)
self.assertIn(2024, all_years)

def test_normalize_and_filter_stream_caret_delimited(self):
"""Verify record normalization and entity filtering for caret-delimited format."""
sample_caret = (
'County Code^District Code^School Code^Type ID^Test Year^Test ID^Student Group ID^Grade^'
'Total Students Tested with Scores^Mean Scale Score^Percentage Standard Exceeded^'
'Percentage Standard Met^Percentage Standard Met and Above^Percentage Standard Nearly Met^Percentage Standard Not Met\n'
'00^00000^0000000^4^2024^1^1^3^100^2400.0^20.0^25.0^45.0^25.0^30.0\n'
'01^00000^0000000^5^2024^1^1^3^60^2410.0^22.0^28.0^50.0^20.0^30.0\n'
'01^12345^6789012^7^2024^1^1^3^50^2350.0^10.0^20.0^30.0^30.0^40.0\n'
)
stream = io.StringIO(sample_caret)
rows = list(
download.normalize_and_filter_stream(stream,
keep_all_entities=False))
# School-level entity (Type ID 7) should be filtered out; State (4) and County (5) kept
self.assertEqual(len(rows), 2)
# State record
self.assertEqual(rows[0][0], '00')
self.assertEqual(rows[0][1], '2024')
self.assertEqual(rows[0][2], '1') # Student Group ID
self.assertEqual(rows[0][3], '3') # Grade
self.assertEqual(rows[0][4], '1') # Test ID
self.assertEqual(rows[0][5], '100') # Tested with scores
# County record
self.assertEqual(rows[1][0], '01')
self.assertEqual(rows[1][1], '2024')

def test_normalize_and_filter_stream_comma_delimited_legacy_headers(self):
"""Verify record normalization for comma-delimited data with older header aliases."""
sample_comma = (
'"County Code","District Code","School Code","Type ID","Test Year","Test Id","Subgroup ID","Grade",'
'"Students with Scores","Mean Scale Score","Percentage Standard Exceeded",'
'"Percentage Standard Met","Percentage Standard Met and Above","Percentage Standard Nearly Met","Percentage Standard Not Met"\n'
'"00","00000","0000000","4","2016","1","1","4","120","2450.0","25.0","30.0","55.0","20.0","25.0"\n'
)
stream = io.StringIO(sample_comma)
rows = list(
download.normalize_and_filter_stream(stream,
keep_all_entities=False))
self.assertEqual(len(rows), 1)
self.assertEqual(rows[0][0], '00')
self.assertEqual(rows[0][1], '2016')
self.assertEqual(rows[0][2], '1') # Mapped Subgroup ID -> index 2
self.assertEqual(rows[0][3], '4') # Grade
self.assertEqual(rows[0][4], '1') # Mapped Test Id -> index 4
self.assertEqual(rows[0][5],
'120') # Mapped Students with Scores -> index 5

def test_normalize_and_filter_stream_keep_all_entities(self):
"""Verify keep_all_entities=True retains school and district level entities."""
sample_caret = (
'County Code^District Code^School Code^Type ID^Test Year^Test ID^Student Group ID^Grade^'
'Total Students Tested with Scores^Mean Scale Score^Percentage Standard Exceeded^'
'Percentage Standard Met^Percentage Standard Met and Above^Percentage Standard Nearly Met^Percentage Standard Not Met\n'
'00^00000^0000000^4^2024^1^1^3^100^2400.0^20.0^25.0^45.0^25.0^30.0\n'
'01^12345^6789012^7^2024^1^1^3^50^2350.0^10.0^20.0^30.0^30.0^40.0\n'
)
stream = io.StringIO(sample_caret)
rows = list(
download.normalize_and_filter_stream(stream,
keep_all_entities=True))
self.assertEqual(len(rows), 2)


if __name__ == '__main__':
unittest.main()
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
key,value
header_rows,1
word_delimiter,""""""
input_delimiter,^
output_columns,"observationDate,observationAbout,variableMeasured,value,unit,scalingFactor"
Loading
Loading