Projects for Fall 2026
Project 1: Civic
Contact Person: Alex (me)
Preferred Contact Method: Email, Videoconference (e.g., Zoom, Google Hangouts), Slack
Description of Organization
Our platform, Revere, allows elected representatives to better interface with their constituents. We support constituent message intake, understanding, and representative outreach. The platform is live, currently supporting 250k constituents, and is in the final approval cycle for the US House of Representatives.
Project Description
Congress received 81 Million messages last year, 70% of which were form (pre-written) letters. Many of these letters also had personal messages added to the forms. Offices of representatives want to know these personal messages, and current systems are generally not capable of cleanly separating the portion of the message that a person wrote from the form letter.
You’ll be working on a public dataset to develop a methodology to separate the portion of these letters that were written by the constituent from the form letter, then perform structured extraction over the written letter portion.More info
-
How many rows of data do you imagine will be involved in total?:
- hundreds of thousands
-
What kind of data will the project involve?:
- non-rectangular (e.g., JSON, XML, etc.), We can provide pre-formatting if needed.
-
Tags:
- Statistical modeling, Predictive modeling (aka, machine learning), Data analysis, Data visualization, Text analysis / regular expressions / natural language processing, Artificial Intelligence
-
Is there anything else you’d like to share with the students?:
- You’ll have the opportunity to work with Civic AI scientists and engineers to develop the methodology. We look forward to working with you!
Repository
https://github.com/sds-capstone/civic-f26Video
https://drive.google.com/file/d/1YEB6z7Lv7ZvpJaj40-k-OfI32lj9i8So/view?usp=drive_link
Project 2: The Connecticut Mirror
Contact Person: Angela Eichhorst
Preferred Contact Method: Email, Videoconference (e.g., Zoom, Google Hangouts)
Description of Organization
The Connecticut Mirror is a local nonprofit, digital only newsroom that reports on public policy, government and politics. We produce original, in-depth, non-partisan journalism that informs Connecticut residents about the impact of public policy, holds government accountable, and engages and amplifies diverse voices and perspectives. Since its founding in 2010 the Mirror has grown to a staff of 30 and serves over 300,000 readers a month.
Project Description
Journalists want to include cross-Connecticut comparisons and historical trend data in their public policy impact stories, but don’t know how to use US Census data to their advantage. We are looking for an easy to use dashboard/app/website that displays potential census questions by topic, presents visualizations of historical and town-by-town trends, and potentially allows for the export of data visualizations to news articles. At the town level, many census questions have an extremely high margin of error that makes them inappropriate to use for analysis, so the tool could also warn against or prevent the display of data where the margin of error is too high. An AI chatbot could assist journalists in discovering what Census questions are available for their topic. This project can be done in Python, R, or any other advantageous languages.
More info
-
How many rows of data do you imagine will be involved in total?:
- millions
-
What kind of data will the project involve?:
- rectangular (i.e., rows and columns, spreadsheets, CSVs), SQL database
-
Tags:
- Web app development, Web scraping, Data analysis, Data wrangling, Data visualization, Geospatial analysis (i.e., shapefiles, mapping), Artificial Intelligence
-
Is there anything else you’d like to share with the students?:
- NA
Repository
https://github.com/sds-capstone/ctmirror-f26Video
https://drive.google.com/file/d/1DYtEizWAiiaRy_CIieBkUAMlk4PhcWcg/view?usp=drive_link
Project 3: Dance Data Project
Contact Person: Catherine Spratt
Preferred Contact Method: Email, Videoconference (e.g., Zoom, Google Hangouts)
Description of Organization
Dance Data Project® promotes gender equity in the dance industry, including but not limited to ballet companies, by providing a metrics-based analysis.
Through our research, programming, resources, and advocacy, DDP showcases and uplifts women throughout the dance industry. We focus on leaders, both artistic & administrative, and artists of merit: choreographers, photographers, lighting, costume, and set designers, commissioned composers, film directors/producers, etc.Project Description
Dance Data Project®, the organization bringing rigorous, metrics-based analysis to gender equity in dance, is launching a new research initiative to collect and analyze LinkedIn profiles of Artistic Directors across the Largest 150 Ballet Companies. The project will test the hypothesis that women must accumulate more leadership experience, more advanced education, and be older than their male counterparts before earning the title of Artistic Director. The project builds directly on the foundational work of Professor Herrera-Guzmán and Professor Grundstrom, who identified a critical gap in the field. As Grundstrom writes, “What is missing, however, is the narrative surrounding the female leadership experience” (Grundstrom 24). By quantifying that leadership pipeline for the first time, this research, like all of DDP’s 75+ publicly available studies, has real potential to be cited by scholars and journalists alike. It will be published under the Dance Data Project® name with Smith College student researchers credited as authors. This will be the fifth round of our collaboration with a Smith College Data Science Capstone team, following the four previous partnerships with research including the Cost of Living Interactive Maps (Fall 2024) and the Endowments and Book Building Report 2024 (September 2024). The Cost of Living Interactive Maps illustrate the financial disparity between the average cost of living in a state in comparison to the salaries of Artistic and Executive Directors in the state, and the Endowments and Book Building Report 2024 examined the financial foundations of ballet companies, spanning from fiscal year 2016 to fiscal year 2023. Those projects demonstrate the success of a rewarding collaboration between Smith College and Dance Data Project®, and the upcoming research initiative offers more students a similar experience. Do you want to make a lasting impact on an entire economic sector — one that notoriously bars women from leadership roles? Now’s your chance.
More info
-
How many rows of data do you imagine will be involved in total?:
- hundreds
-
What kind of data will the project involve?:
- rectangular (i.e., rows and columns, spreadsheets, CSVs)
-
Tags:
- Statistical modeling, Predictive modeling (aka, machine learning), Web scraping, Data analysis, Data visualization, Text analysis / regular expressions / natural language processing, Artificial Intelligence
-
Is there anything else you’d like to share with the students?:
- NA
Repository
https://github.com/sds-capstone/dancedata-f26Video
https://drive.google.com/file/d/1XCWMw9WJA6nFTc2wjhu4alKA4vx-J3g_/view?usp=sharing
Project 4: Office of the District of Columbia Deputy Mayor for Education
Contact Person: `Abdu’l-Karim Ewing-Boyd
Preferred Contact Method: Email, Microsoft Teams
Description of Organization
The Office of the District of Columbia Deputy Mayor for Education (DME) is responsible for developing and implementing the Mayor’s vision for academic excellence and creating a high-quality education continuum from early childhood to K-12 to post-secondary and the workforce, literally from cradle to career. The three major functions of the DME are overseeing a District-wide education strategy, managing interagency and cross-sector coordination and providing oversight and support for the District’s education and workforce agencies.
Project Description
I am trying to solve poverty, human development, and good governance. For this piece of that work, I am looking to replicate and specify the Urban Institute’s Upward Mobility Data Dashboard at the local level to inform data driven, equity focused policy. DC is woefully economically extreme, with the richest parts of town masking deep impacts from intergenerational poverty in the citywide data available through Urban’s Dashboard. I imagine a final product that is very similar to the current Urban Institute dashboard at the tract, block, or neighborhood level.
More info
-
How many rows of data do you imagine will be involved in total?:
- tens of thousands
-
What kind of data will the project involve?:
- rectangular (i.e., rows and columns, spreadsheets, CSVs), non-rectangular (e.g., JSON, XML, etc.), SQL database
-
Tags:
- Statistical modeling, Web app development, Data analysis, Record linkage, Data wrangling, Data visualization, Geospatial analysis (i.e., shapefiles, mapping)
-
Is there anything else you’d like to share with the students?:
- This is an opportunity to work on a project that will be used immediately for real, government led, place-based initiatives to increase equity and justice.
Repository
https://github.com/sds-capstone/dced-f26Video
https://drive.google.com/file/d/1Wk4AS0PTvu6AJG_yW0zKihcpd8pEO3Ia/view?usp=drive_link
Project 5: redacted
Contact Person: Mariel Finucane
Preferred Contact Method: Email, Videoconference (e.g., Zoom, Google Hangouts)
Description of Organization
We’re a large, not-for-profit health plan serving 3 million members in Massachusetts and across the country.
Project Description
Problem: The U.S. Center for Medicare and Medicaid Services (CMS) created a Star Rating system to help Medicare beneficiaries compare plans on quality measures like controlling blood pressure or completing colorectal cancer screenings. The Star Rating system also incentivizes plans to deliver high-quality care by rewarding top performers with bonus payments. A plan’s Star Rating on each quality measure depends on where its score falls relative to CMS-defined cutpoints (e.g., “you need at least a 78% screening rate to earn 4 stars on this measure”). The challenge is that cutpoints are not fixed — CMS recalculates them each year based on how all plans performed, meaning they typically rise as the industry improves. Critically, cutpoints are set after the year ends: a plan performing right now in 2026 won’t learn what cutpoints its scores will be graded against until after the year is over. Without a way to project where cutpoints are headed, plans are essentially flying blind when setting performance targets.
Goal/Objective: The goal of this project is to build a model that projects Medicare Stars cutpoints before they are officially released by CMS. Students will explore two complementary approaches: (1) projecting cutpoints directly based on historical cutpoint data, and (2) projecting them indirectly by forecasting measure scores for all plans nationally and simulating where the clustering algorithm would set thresholds. The project will compare the accuracy of these two approaches.
Final Product: A final product will take the form of a write-up or dashboard comparing the performance of the two modeling approaches, and a dataset of best-performing cutpoint projections for an upcoming Star Ratings year.
Why it matters: Because achieving four or more stars unlocks significant CMS bonus payments, accurate cutpoint projections are a high-stakes input to strategic planning — a single cutpoint boundary can mean millions of dollars in revenue for a health plan. But there’s a mission-driven dimension too: plans that can better anticipate where cutpoints are headed can focus quality improvement efforts more effectively, ultimately delivering better care to Medicare beneficiaries. Improving our ability to project cutpoints is a step toward a health system where plans compete not just on price, but on genuinely raising the bar for patient outcomes.More info
-
How many rows of data do you imagine will be involved in total?:
- hundreds of thousands
-
What kind of data will the project involve?:
- rectangular (i.e., rows and columns, spreadsheets, CSVs)
-
Tags:
- Statistical modeling, Predictive modeling (aka, machine learning), Web scraping, Data analysis, Data wrangling, Data visualization
-
Is there anything else you’d like to share with the students?:
- Thanks Ben!
Repository
https://github.com/sds-capstone/hcstars-f26Video
Project 6: 99P Labs / Honda Research Institute
Contact Person: Ryan Lingo
Preferred Contact Method: Email, Videoconference (e.g., Zoom, Google Hangouts), Slack, GitHub
Description of Organization
99P Labs is Honda Research Institute’s innovation lab for exploring future mobility, energy, software, AI, and human-centered technology.
Project Description
This research project invites data science students to explore how human judgment matters as AI takes on more of the work of writing code. Students will study how people read, understand, evaluate, and trust code in AI-assisted settings, while developing open-ended ways to measure those skills in practice.
More info
-
How many rows of data do you imagine will be involved in total?:
- thousands
-
What kind of data will the project involve?:
- Don’t know
-
Tags:
- Statistical modeling, Predictive modeling (aka, machine learning), Web app development, Data analysis, Data visualization, Experimental design, Text analysis / regular expressions / natural language processing, Software development, Artificial Intelligence
-
Is there anything else you’d like to share with the students?:
- NA