Data mining with historical documents

The last seminar held by the Vision and Graphics Laboratory was about data mining with historical documents. Marcelo Ribeiro, a master student at the Applied Mathematics School of the Getúlio Vargas Foundation (EMAp/FGV), presented the results obtained with the application of topic modeling and natural language processing on the analysis of historical documents. This work was previously presented at the first International Digital Humanities Conference held in Brazil (HDRIO2018) and had Renato Rocha Souza (professor and researcher at EMAp/FGV) and Alexandre Moreli (professor and researcher at USP) as co-authors.

The database used is part of the CPDOC-FGV collection and essentially comprises historical documents from the 1970s belonging to Antonio Azeredo da Silveira, former Minister of Foreign Affairs of Brazil.

The documents:

Dimensions
• +10 thousand documents
• +66 thousand pages
• +14 million tokens / words (dictionaries or not)
• 5 languages, mainly Portuguese

Formats
• Physical documents
• Images (.tif and .jpg)
• Texts (.txt)

The presentation addressed the steps of the project, from document digitalization to Integration of results into the History-Lab platform.

The images below refer to the explanation of the OCR (Optical Character Recognition) phase and the topic modeling phase:

Presentation slides (in pt) can be accessed here. This initiative integrates the History Lab project, organized by Columbia University, which uses data science methods to investigate history.