Tech

City Libraries Pilot AI Transcription to Preserve Immigrant Oral Histories

Public librarians are testing speech-to-text models tuned to multilingual accents to digitize decades of community interviews while wrestling with accuracy and consent.

By Caroline O'Neill · March 13, 2026 · 4 min read

City Libraries Pilot AI Transcription to Preserve Immigrant Oral Histories

In a test that could reshape how New Yorkers remember their neighborhoods, city libraries across Queens and Brooklyn have begun piloting artificial-intelligence transcription tools to digitize decades of immigrant oral histories recorded on aging tapes and digital files, a push that librarians say is urgent as fragile cassettes deteriorate in basements on Northern Boulevard and in community rooms under the Jackson Heights branch on 37th Avenue.

The pilot, quietly launched this winter, pairs librarians with linguists and a local startup to run speech-to-text models tuned to South Asian, Caribbean, Mandarin, Cantonese and Spanish-influenced English — a response, librarians say, to the ways phonetics and code-switching have repeatedly failed generic transcription software.

“We are racing against time and leaky basements,” said Marisol Reyes, head of the oral history program at Elmhurst Community Library, who oversees the project across eight neighborhood branches. “If we can make interviews searchable and readable without losing the honor of the speaker, we give this city a chance to see itself more whole.”

The technology in the pilot was developed in a partnership between City Libraries and PolyLingua Labs, a Ridgewood startup, alongside the Metropolitan Applied Language Lab at Queensborough University; the model was fed anonymized samples supplied by neighborhood groups and refined with lexicons developed from volunteer transcribers in Sunset Park and Flushing. The software emphasizes speaker diarization — tagging who is speaking when — to help separate overlapping conversation in crowded living-room interviews recorded at community centers on Jamaica Avenue.

Numbers underscore both the scale and the risk: the pilot currently covers 12 branches, processes roughly 18,700 interview files totaling 2,400 recorded hours, and handles material in at least seven languages and a dozen dialects. Initial baseline testing put out-of-the-box transcription accuracy at about 66 percent for heavily accented English and multilingual passages; after tuning with local datasets, the pilot reports an average word-error rate improvement to 86 percent, a reduction in manual correction time from an estimated 10 minutes per hour of audio to about 4, and a projected storage need of 22 terabytes for processed and raw files combined.

“A misheard name erases a person from the story,” said Asha Patel, senior archivist for the South Asian Oral History Project, which contributed more than 3,000 interviews to the pilot. “We’re not just fixing grammar; we’re protecting genealogy, labor histories and migration trajectories. When the software mislabels an employer or hometown, we have a responsibility to correct it.”

That responsibility is central to the pilot’s consent and privacy framework. City Libraries says every collection partner has been asked to re-consent interviewees where possible and to adopt tiered access so that sensitive passages can be redacted from public view. “We built consent into the workflow,” said Daniel Morales, deputy director for City Libraries, who manages digital services. “Transcription is not an automatic publish button. Nothing goes online without review, and communities set the guardrails.”

On a weekday morning in Jackson Heights, a volunteer tucked into a sunlit alcove at the Woodside Branch listens to a 1990s cassette of South Asian garment workers while correcting the AI’s rendering of a village name; across town in Sunset Park, an organizer named Nguyen Tran reviews interviews about the 1980s Chinese immigrant restaurant scene, praising the speed but flagging errors where colloquialisms were flattened. “The machine gets the structure, but not the soul,” Tran said, adding that volunteers still spend hours restoring local color and proper nouns. “If you don’t have people in the loop, you lose the neighborhood.”

Experts working on the project say algorithmic bias remains the main technical obstacle. “Models trained on broadcast news or single-accent corpora systematically underperform on code-switched, community-recorded speech,” said Prof. Leila Haddad, a computational linguist at the Metropolitan Applied Language Lab who helped audit early model outputs. “We’ve seen systematic errors with names, place markers and occupational terms. Human-in-the-loop review and transparent error reporting are essential to avoid creating a secondary archive of mistakes.”

For now, the pilot will run for 12 months with quarterly public audits and a community advisory board drawn from partner organizations in Elmhurst, Ridgewood, Flushing and Bay Ridge; funders include a municipal cultural preservation grant and private philanthropy, and libraries say any public release will be accompanied by searchable annotations and version histories so users can see original audio, the raw AI transcript and the final edited record. If outcomes align with community expectations on accuracy and consent, City Libraries officials say they intend to scale the program to more borough branches next year, while continuing to refine the models and expand volunteer training — a cautious step forward that, organizers hope, will preserve messy, human memory without erasing the voices that made the city what it is.