AI, OCR, and Machine Translation for Low-Resource Languages
Project Description
Low-resource languages often lack the large datasets and digital tools available for widely used languages. This project explores how recent advances in artificial intelligence can support the digitization and translation of low-resource historical languages, with a particular focus on Manchu.
Students will work with OCR and machine translation workflows for Manchu historical sources and examine how different AI models perform when training data and digital resources are limited. The project combines digital humanities, historical research, and applied AI.
Students will work with OCR and machine translation workflows for Manchu historical sources and examine how different AI models perform when training data and digital resources are limited. The project combines digital humanities, historical research, and applied AI.
Supervisor
CHUNG, Yan Hon Michael
Quota
2
Course type
UROP1100
UROP2100
Applicant's Roles
The project will focus on 2 related tasks: converting scanned historical documents into machine-readable text through OCR, and using AI models to assist with translation between Manchu and Chinese or English.
Depending on their interests and background, students may work on preparing and correcting OCR data, comparing outputs from different AI models, evaluating recognition and translation errors, constructing bilingual datasets, or testing methods for improving model performance. Students may also explore how these technologies can be applied to other low-resource languages.
No prior knowledge of Manchu is required. Basic training in the language, historical sources, and relevant digital tools will be provided. Programming experience is helpful but not required.
Depending on their interests and background, students may work on preparing and correcting OCR data, comparing outputs from different AI models, evaluating recognition and translation errors, constructing bilingual datasets, or testing methods for improving model performance. Students may also explore how these technologies can be applied to other low-resource languages.
No prior knowledge of Manchu is required. Basic training in the language, historical sources, and relevant digital tools will be provided. Programming experience is helpful but not required.
Applicant's Learning Objectives
Applicants may be involved in:
Preparing and organizing digitized historical sources.
Checking and correcting OCR-generated texts.
Creating and verifying Manchu-Chinese or Manchu-English parallel texts.
Testing OCR, vision-language, and machine translation models.
Comparing model outputs and identifying recurring errors.
Assisting with the construction and documentation of research datasets.
Conducting basic quantitative or qualitative analysis of model performance.
Presenting progress and findings to the research team.
Preparing and organizing digitized historical sources.
Checking and correcting OCR-generated texts.
Creating and verifying Manchu-Chinese or Manchu-English parallel texts.
Testing OCR, vision-language, and machine translation models.
Comparing model outputs and identifying recurring errors.
Assisting with the construction and documentation of research datasets.
Conducting basic quantitative or qualitative analysis of model performance.
Presenting progress and findings to the research team.
Complexity of the project
Moderate