ParamergeParamerge

Domain Atlas

Child welfare & family services

Predictive screening and profiling where the cost of both false alarms and misses lands on families — and where the human override layer has measurably mattered.

Use cases

What AI is doing here

Maltreatment call screening

Predictive

Risk scores supporting screen-in/screen-out decisions on child-maltreatment referrals.

Family risk prediction

Predictive

Longitudinal risk models over family and administrative data to prioritize investigation or services.

Early-help profiling

Predictive

Mining council/agency data to flag families for preventive outreach before crisis.

Case notes as training data

Predictive

Using narrative case records to train predictive models — importing the biases and errors those records contain.

Case files

What has gone wrong and right

Documented deployments, presented as model organizations calibrated to the evidence, with full citations.

The score nobody sees: New York City's concealed severe-harm QA algorithm

New York City, New York, United States

Since May 2018, New York City's child-welfare agency has scored every open child-protection investigation at day 10 with an in-house machine-learning model, rank-ordering cases by predicted likelihood of substantiated physical or sexual abuse within 24 months to fill a capacity-limited 'quality assurance' review worklist — about 3,000 of roughly 50,000 investigations a year. Nobody who acts on the output sees it: not families, not their attorneys, not frontline caseworkers, and — by the city register's own words — not even the QA reviewers working the flagged list. The agency's own internal audit conceded the training data likely carried systemic biases, that geography may act as a partial race proxy, and that flag predictions are 'more likely to be incorrect than correct'; the agency told state auditors the model outperforms experienced caseworkers (a self-claim with no public methodology) and that there is 'no basis for a complaint.' The oversight graph is blocked-but-not-inert: a 2023 state audit produced promises, the charter-mandated inspector general is statutorily severed from the records — full access in only 1 of 18 child fatalities with prior agency involvement in 2025 — and the algorithmic-accountability office the Council legislated in November 2025 remains unverified as operational, while the model expanded in 2025 into the very record class state law prohibits the inspector general from touching.

The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments

Colorado, USA (statewide; state-supervised, county-administered system across 64 counties; audit loop through the Office of the Colorado Child Protection Ombudsman, an independent judicial-department agency, and the Colorado General Assembly)

Colorado runs every screened-in child welfare referral through two statewide instruments: a structured Family Safety Assessment (10 danger items gated by 5 judgment-heavy threshold criteria) and an actuarial Family Risk Assessment scored largely from prior system contact stored in the Trails record system, used by 64 county-administered agencies under one state rule. In 2024 the legislature handed the state's independent Child Protection Ombudsman — not the child welfare agency — a funded mandate (HB 24-1046, $109,392) to procure a third-party audit of the tool suite against nine statutory criteria. ICF's 176-page audit (released March 2, 2026; 50 recommendations) found the tools well-aligned with policy on paper but inconsistently implemented: 68% of caseworkers cannot complete assessments in real time, vignette agreement collapses to 59–63% on neglect and poverty scenarios, the risk tool's weighting of old reports 'disproportionately impacts minority families and perpetuates systemic bias' (with any past DV incident coded as recent), and race and ethnicity are documented so inconsistently that the disparity analysis the statute requested could not be run. The audit recommends revising or replacing the actuarial tool — and as of July 2026 the oversight loop has produced statute, budget, and public information, but no tool change.

The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to

Florida, Georgia, Virginia (attempted Tennessee), United States

Since 2018, Adoption-Share's Family-Match — a proprietary 'relational fit' matching algorithm built by former eharmony researchers — has paired foster children awaiting adoption with self-recruited prospective parents in Florida, Georgia, and Virginia, emitting a ranked candidate-parent list per child that caseworkers then vet toward trial placement and legal adoption. The independently observed results are stark: Virginia's two-year test produced one known adoption (an official's statement to the AP; the pilot's end is undated in the record, circa 2020) and Georgia's year-long pilot produced two before ending in October 2022 — yet both states resumed business with the vendor, Georgia for free in July 2023 after direct lobbying, Virginia via a larger contract for a rebranded recruitment portal ($212,546 spent in SFY2023). In Florida, the showcase state, the vendor's own ledger claimed 603 placements and 431 adoptions over five years — figures its partner agencies could not verify: one agency's records showed 76 tool-attributed placements with no documented adoption, another counted 8 adoptions from 22 matches while making hundreds of both without the tool. The AP's November 2023 investigation (Sally Ho and Garance Burke) found the vendor owned the outcome data, refused states access to the algorithm even on request, and — per its own April 2023 user guide — had caseworkers document matches made outside the tool inside the system. Only Tennessee interrogated the data schema before deployment; only Tennessee never deployed.

The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record

District of Columbia, United States

In June 2025, DC's Child and Family Services Agency shipped the child-welfare domain's first documented policy-retriever deployment: CORA, a staff-facing chatbot embedded in STAAND, the agency's new case-management system, answering policy and procedure questions from a small expert-validated corpus — with a published 16-page pre-deployment governance report (one of only three district-wide) whose centerpiece guardrail is structural: the chatbot 'does not draw from confidential case files' and has no access to STAAND case data. Within six months the agency's own tip sheets documented the guardrail's blind side: a CORA phone app that ingests case documents, photos of handwritten notes, and voice dictation, and saves AI-drafted contact notes into the STAAND case record — the write direction the no-case-files rule never governed, superseding the May 2025 report's claim that generated text was 'guidance, not case documentation.' The same report conceded per-output human validation is impossible; verification is batch-mode (expert curation before, monthly internal accuracy reports after, an annual corpus review against monthly platform builds), and an in-product warning — 'answers may not be fully accurate. Consult your supervisor as needed' — delegates per-query vigilance to staff. The whole loop sits inside one agency, in its first era without a court monitor in three decades; no independent audit or published accuracy data exists, and every outcome figure is a vendor claim from a customer story that never names CORA.

System map

Who is in the system and what pushes on it

Who is in the system

  • Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
  • Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
  • Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
  • Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
  • Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
  • Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.

Dominant pressures

  • Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover. Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

Governance

Questions leaders should be asking

  1. 1. What does a risk score change about a worker's next action — and is that mapping written down anywhere?
  2. 2. Are overrides tracked, and does anyone know whether they are improving or degrading equity?
  3. 3. If the tool were saturating workers with alerts, how would leadership find out?
  4. 4. What would trigger discontinuation, and who holds the authority to trigger it?
  5. 5. Who receives a report that the tool has harmed someone — a family, a worker, an advocate — on what clock, and what has to happen next once they have it?

For the actions behind these questions, see the Practice Library.

Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.

Work With Paramerge

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

ConceptualThe volume's governance chapter places two standing duties on an agency that deploys AI, both distinct from an…

The volume's governance chapter places two standing duties on an agency that deploys AI, both distinct from any pre-deployment approval. First, an incident-reporting protocol that enables timely identification and remediation of algorithmic harm - discriminatory treatment, a biased risk assessment, a misdiagnosis, a breach of confidentiality. Second, transparent channels through which both the people served and the practitioners can report concerns or unexpected effects, so the accountability loop closes after deployment rather than ending at approval. The chapter is a conceptual synthesis and is cited as one: its four-tier social-work risk taxonomy is labeled by its own author as an original construction, informed by but not derived from binding regulation. It may be cited as a framework and must never be presented as a regulatory classification of any deployment in this registry.

huang2026bAcademicSave

Huang, J., Yang, F., & Lee, J. (2026). Ethical Challenges and AI Governance in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_22

doi.org/10.1007/978-3-032-18443-6_22

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-ethics, ai-governance, social-work