Tohono O’odham FLEx Dictionary

Published:

Introduction

The main focus of this project was to digitize a dictionary of the O’odham language, spoken in the Sonoran Desert of Arizona and northern Mexico by the Tohono O’odham (“Desert People”). The source text is in a flat, text-based format, and my goal was to transform it into a structured representation that could be more easily analyzed and useful to the Tohono O’odham and linguists. While the original dictionary contains rich lexical and grammatical information, its lack of consistent structure makes it difficult to use for both linguistic research and computational applications.

Background

Tohono O’odham is an Indigenous language spoken in the southwestern United States and northern Mexico, with significant cultural and linguistic importance. The data contained in this dictionary was originally gathered by Madeleine Mathiot over two years (1958-1960), while staying at the Covered Wells Mission in Arizona. Much of the data comes directly from one Tohono O’odham individual, Jose Pancho. Her original dictionary1 was published in the early ’70s and was the most comprehensive dictionary of the Tohono O’odham language. In the late 90s and early 2000s, Dr. Michael Hammond and Dr. Ofelia Zepeda (along with several students) began work on putting the dictionary on the internet. The typed pages were scanned with OCR and errors were manually fixed. The dictionary website created by Dr. Hammond and his team is no longer supported by the university, making it no longer usable. The scanned dictionary pages were the starting point for my project. The aim of this project was to get all of the data into FLEx (FieldWorks Language Explorer), which is widely used for lexical database management and language documentation.

Problem Statement

The primary challenge of this project lies in converting semi-structured dictionary text into a fully structured format. Having the data in a structured database would allow for more complex searches in the dictionary, such as isolating searches to certain fields. The source data contains inconsistencies, such as varying use of punctuation, abbreviations, and embedded notes. Additionally, certain patterns—such as “see” references or parenthetical notes—can be difficult to distinguish from core lexical information. These issues make it difficult to reliably extract fields like part of speech, senses, and morphological features using straightforward methods. As a result, a more nuanced parsing approach was required to handle these ambiguities. Below is a snippet from the first page of the dictionary, in the format it was given to me (.docx/.pdf). You can see from this snippet the structure I had to contend with: boldface, page numbers, guide words, and the alphabetical range header.

Here is a UML diagram of the structure for the dictionary entries.

Approach

To address these challenges, I developed a multi-step parsing pipeline that processed each dictionary entry individually. The approach begins by identifying entry boundaries and then progressively extracting specific components such as headwords, parts of speech, and definitions. A key design decision was to represent each entry as a Lexeme object, allowing for a clear and extensible data structure. Regular expressions and rule-based heuristics were used to identify patterns within the text. For initial testing, a carefully selected test set of data was used to ensure that all the different entry types were being correctly handled.

Implementation

The parser was implemented in Python, using built-in libraries such as re for regular expression pattern matching. Each dictionary entry is processed and converted into a structured object with predefined fields. One of the main challenges was correctly handling nested or overlapping annotations, such as parentheses containing additional metadata. Another difficulty involved distinguishing between true lexical content and cross-references introduced by terms like “see.” To address these issues, the parser contains special functions that categorize different types of information. The resulting data structure allows for easier manipulation and export to other formats. To import the data into FLEx, I reorganized the data into a Standard Format Marker (SFM) file. This format provided a practical way to preserve the structure of the parsed data while allowing it to be imported into FLEx.

Results

The final output is a usable dictionary in the FLEx dictionary software and a downloadable dictionary app for Android users. The parser was far from perfect, but in many cases the parser successfully extracted key components such as headwords, parts of speech, and definitions. An originally loosely structured entry was transformed into a format where each element is explicitly labeled and accessible. However, some edge cases remained challenging and required additional refinement. Even after the final import to FLEx, there was still a large amount of manual work due to the innumerable inconsistencies in the entries, plus additional nuances and details that were not taken into account when originally writing the parser. Below are several examples of how the parser correctly and incorrectly handled various entries.

Example 1

Here is an example of a typically structured dictionary entry that was correctly parsed, with all the necessary items being mapped to the correct fields in FLEx. The headword maps to Lexeme Form, definition to definition, Part of speech to Grammatical Info, and the examples which map to a field for the sentence in the vernacular (Tohono O’odham), and the translation (English).

Example 2

Here is an example of an entry where the headword is actually part of a phrase, but the phrase itself is not a part of the lexicon.

This was a structure that I did not take into account when I was writing the parser, and therefore did not get parsed correctly. The parser recognized “‘o’ohon in Jioş-‘O’ohon” as the entire headword.

I fixed entries like this manually, by putting “in Jioş-‘O’ohon” in the Restrictions field, to show that this particular usage of ‘o’ohon is restricted to its use in the whole phrase “Jioş-‘O’ohon.”

Example 3

Here is another example of a common incorrect parsing. In the below image, the headword is ‘i’ihugga, followed by a cross reference to another entry vuḑ ‘ihugga. My parser should be easily able to handle this entry because it fits the standard format of headword “see” POS Cross-referenced word “=” definition. The problem with this particular case is that the word order of vuḑ ‘ihugga is not in the same format that the actual entry for that lexical item is in.

Below is how vuḑ ‘ihugga actually appears in the dictionary. It is entered as ‘ihugga vuḑ, which the parser did not recognize to be the same thing as vuḑ ‘ihugga.

This is how this was parsed in FLEx. This is another case where I manually fixed entries of this variety.

Here is how the final entry looked after I manually fixed the cross-reference.

Bonus

I also turned this project into an Android application, using the Dictionary App Builder software by SIL. This is available here for download, and I hope to have it available soon on the Google Play Store. This app will hopefully serve as a useful tool for anyone needing to access Tohono O’odham language data. Here are some screenshots from the application. There are instructions for how to download the app on your Android device below the screenshots.

Download instructions

How to install:

  1. Download .apk file
  2. Open file
  3. Select “Allow unknown apps”
  4. Install the app

Conclusion

One of the most challenging aspects of this project was dealing with inconsistencies in the source data. Small variations in formatting often required additional rules or exceptions in the parser. This highlighted the difficulty of working with real-world linguistic data, which rarely conforms to strict standards. At the same time, the project provided valuable insight into the structure and complexity of dictionary entries. It also reinforced the importance of balancing precision with flexibility when designing parsing systems.

In summary, this project demonstrates a method for converting unstructured dictionary data into a structured format suitable for computational use. By developing a custom parser, it is possible to extract meaningful linguistic information from complex text. This work contributes to broader efforts in language documentation and the development of tools for under-resourced languages. Ultimately, structuring this data makes it more accessible for both linguistic analysis and practical applications, and is hopefully a useful tool for the Tohono O’odham people.

Future Work

Future improvements could focus on increasing the robustness of the parser, particularly in handling edge cases. Additionally, a user interface could be developed to allow for easier browsing and searching of the dictionary. More advanced techniques, such as machine learning, could also be explored to improve parsing accuracy. There remain a few issues in the dictionary entries themselves, as I was not able to look at every single entry to ensure that everything parsed 100% correctly. I plan to do a more comprehensive review in the future and make an update to the app.

Click here.

  1. Mathiot, Madeleine. A Dictionary of Papago Usage. Indiana University, 1973. ↩