There are many ways in which we believe ArchiHUB can help with document transcription. From the very beginning, one of its core objectives has been to bring information into text, whether through audio transcription, image description, or good old OCR.
Today, this can be done with relative ease for contemporary documents. But what happens when we turn our attention to historical archives?
Something interesting happens here. Do you remember the old CAPTCHAs that asked users to identify distorted letters and words? It turns out that many of those systems were doing more than preventing spam. They were generating training data to improve the transcription of historical documents.
Historical documents from where, though?


That question is where this project begins.
What if we could replicate that idea and apply it to Colombian archives? What if the simple act of solving a challenge could help build the datasets needed to improve the transcription of historical manuscripts, newspapers, letters, and records from across the country?
To make this possible, we developed a Telegram bot dedicated to the task, but with an additional twist: credit where credit is due.
Contributors matter. Acknowledgment matters.
Every user who participates in the transcription process and whose contributions are later validated through consensus will be credited as a collaborator in the dataset. The goal is not only to build a valuable resource, but also to recognize the collective effort required to create it.
Once completed, we plan to publish the dataset openly on Hugging Face, together with the appropriate acknowledgments for everyone who contributed. By making the dataset freely available, we hope to encourage experimentation, support the development of local expertise, and promote the training of models using historical documents and data from our own context.