OTTOMAN TURKISH LINGUISTIC ANALYSIS TOOL
3.0Beta
ENTR
LexiQamus 3.0

Dictionary Digitization and Linguistic Data Model

From source images to analytical and searchable dictionary data: a stage-by-stage account of the method we developed in BuildLQ.

Abstract

We began work on LexiQamus in 2015 and released LexiQamus 1.0, the first version of the platform, on 7 June 2016. We then developed a separate Windows-based application and used it to digitize the full text of Sir James Redhouse's 1890 Turkish and English Lexicon for the first time in February 2020. With this publication, we entered the second major phase of LexiQamus, namely LexiQamus 2.0. Over time, we found that this desktop application was not well suited to working on multiple dictionaries simultaneously, did not provide a sufficiently flexible and granular data model, and that its crowdsourcing capabilities did not reach the level of detail we sought. We therefore decided to develop a more comprehensive, web-based application. We named this application BuildLQ, we are still developing it, and we have been entering data into the system regularly since 2023.

The current major update, released in September 2026, constitutes LexiQamus 3.0. With this version, we have not only renewed the platform's interface and technical infrastructure, but have also placed at the center of LexiQamus's new architecture the detailed dictionary-digitization and linguistic data model that we have been developing for years through BuildLQ. Our fundamental approach in BuildLQ is not merely to digitize historical dictionaries in full text, but to analyze the structure of the dictionary page in as much detail as possible and record every word, word fragment, component, pronunciation, language element, and intra-dictionary relationship separately. We begin by preparing the source images, then process pages and columns, slices, selected lexical units, Ottoman-script values, and their Latin-script counterparts as a sequence of interconnected stages. In this way, the links between stages remain intact as the data moves from the source image down to the smallest analytical unit, and every value remains traceable to the relevant visual region in its source.

At the Type-O stage, we enter Ottoman-script values and analyze them in terms of roots, affixes, components, language, spelling, and other morphological and lexical relationships. At the Type-L stage, we systematically record their Latin-script counterparts, alternative pronunciation possibilities, and spelling-pronunciation relationships. While preserving the physical form of the source text and the dictionary author's own choices, we keep our linguistic analysis separate from them as an independently queryable data layer. In this way, we create not merely a digital copy of the dictionary text, but a relational, extensible, machine-processable dictionary database that makes it possible to establish morphological, semantic, and etymological connections among millions of words across different dictionaries.

We operate this data model together with a multi-layered review and quality-control system based on the roles of User, Reviewer, and Admin. Through blind and cross-review mechanisms, we do not leave data entry solely to individual expertise. Instead, we systematically compare the work of different users and establish shared methodological principles. In these respects, we believe that BuildLQ offers a model applicable not only to the digitization of Ottoman dictionaries but also, more broadly, within the field of digital humanities, particularly in its subfields of digital lexicography and the detailed and relational digitization of historical texts in different languages.

1. Introduction

At first glance, digitizing historical dictionaries may appear to consist of scanning printed text and then converting it into searchable plain text with the aid of Optical Character Recognition (OCR) technologies. Yet the information contained in a dictionary is not limited to the character strings visible on the page. The internal structure of dictionary entries, the continuation of an entry in another column, the distinction between headwords and subheadwords, references made by the author to other entries, word groups not explicitly repeated in the text but recoverable from context, the components of a word, variant spellings, inflected and derived forms, pronunciation information, and linguistic relationships established between words are all part of the dictionary’s information structure.

When we produce only plain text through OCR, these elements are either lost altogether or their relationships to one another cannot be preserved. This problem is particularly acute in Ottoman Turkish dictionaries, where elements from different source languages may combine within a single word, the same word may have numerous spelling and pronunciation variants, and dictionary authors may employ different systems of notation. These characteristics turn digitization into something more than simple text transfer. The BuildLQ system developed within LexiQamus emerged from this need. Our aim is not merely to transfer historical dictionaries into a digital environment, but to record their structural and linguistic information in the most granular form possible and to make the relationships among these components part of the data model itself.

In this article, I describe step by step the dictionary digitization and linguistic data model we have developed within BuildLQ, beginning with the preparation of the source image. I first discuss the development of LexiQamus and the emergence of BuildLQ. I then explain the preparation of source images, the organization of dictionary volumes into projects and sets, the identification of column and entry boundaries, the highlighting of lexical units at the Highlight stage, the entry and analysis of Ottoman-script values at the Type-O stage, and the analytical and categorical recording of Latin-script pronunciations at the Type-L stage. I then present our review and quality-control mechanism, and finally explain why we treat Ottoman-script values as the primary relational layer in the system and how this principle shapes the search architecture.

2. The Development of LexiQamus and the Emergence of BuildLQ

We began work on LexiQamus in 2015. We released LexiQamus 1.0, the first version of the platform, on 7 June 2016. After this initial release, we developed a separate Windows desktop application called LQEdit. Using this application, we published the full text of Sir James Redhouse's celebrated 1890 A Turkish and English Lexicon digitally for the first time in February 2020. Over time, however, we encountered a number of limitations in this Windows-based application. It was not sufficiently flexible and did not allow us to work on multiple dictionaries simultaneously. Its crowdsourcing system functioned, but remained inadequate for our purposes. Nor did it provide the degree of detail and granularity we sought in data entry and linguistic analysis.

After 2020, we therefore decided to develop a new web-based application that would allow us both to process numerous dictionaries within a single system and to analyze dictionary entries in much greater detail. We named this application BuildLQ, and we continue to develop it. Once the system had reached a sufficient level of maturity, however, we began data entry in 2023 and have continued transferring dictionaries into the system ever since.

3. Preparation of Source Images

Before incorporating a dictionary into the system, we prepare its source images. We either scan the dictionary ourselves or obtain an existing digital copy. In both cases, we check for missing or duplicate pages and remove from the working copy any non-textual visual elements that could interfere with later stages. We do not alter the original file, all operations are performed on a second copy, and we keep a record of the changes we make.

When we scan a dictionary ourselves, we pay particular attention to ensuring that light is distributed as evenly as possible across the page surface. Our own scans therefore usually require little intensive image correction afterwards. In digital copies obtained from external sources, by contrast, we more frequently encounter uneven lighting across different parts of a page. In such cases, we apply different Levels adjustments in Photoshop to different regions of the same page in order to standardize the image as far as possible. The result is a legible and workable image in which text and background are clearly separated, which appears almost black and white while retaining grayscale information.

4. Defining Dictionary Volumes as Projects

Once image preparation is complete, we define each volume of the dictionary as a separate project in BuildLQ. If a dictionary consists of more than one volume, we process the volumes as independent projects while also linking them within the system as parts of the same dictionary. In this way, the application treats each volume as a distinct unit of work while preserving the general title of the multi-volume dictionary and the relationship among its volumes. At this stage, we enter the dictionary's title, publication date, author or authors, and general information about the work. We also define the permitted Ottoman- and Latin-script character sets for each project. We consider it especially important to restrict the character set in advance so that only approved characters can be stored in the database. This control eliminates the risk of a word becoming unfindable in search because characters that look identical or nearly identical nevertheless have different Unicode or other encoded values.

This problem is particularly important in Arabic-script text. Characters such as ك, ی, ۃ and many others may look identical or very similar while having different digital values. By controlling the character set at project level, we reduce as far as possible the risk that a word will fail to appear in the database solely for this reason.

5. Units of Work: Sets

For data entry, we divide dictionaries into small units of work, which we call sets. In one of the first dictionaries we processed, Kâmûs-ı Türkî, we initially created five-page sets. Over time, however, we found five-page units tiring for users and therefore reduced the standard set size to three pages in later work. A set thus generally consists of three dictionary pages. In subsequent stages, we successively break these pages down into columns, entries, words, and, where necessary, word components.

6. First Stage: Structural Segmentation of Pages and Columns (Split)

Most historical dictionary pages are organized in columns. At the first stage, which we call Split, we identify the text columns on the page and draw boundaries around them. We do not physically crop the image, instead, we define the coordinates of the columns. We want the areas we draw to encompass the text while excluding non-textual elements such as rules above the column, page numbers, or divider lines between columns. Wherever possible, we use exact rectangular areas. Where this is not possible, we draw polygons that are as close to rectangles as possible. We keep the top and bottom edges perfectly horizontal at zero degrees and align the right and left edges with the slant of the text on the page. We make sure that the boundaries do not touch the text, but we also avoid placing them unnecessarily far away.

Even when a dictionary normally has two columns per page, we may encounter pages with four or more columns, particularly at transitions between letters. In such cases, we do not adhere to a predetermined number of columns. We define as many columns as the actual structure of the page requires.

7. Second Stage: Slice and the Identification of Entry Boundaries

After defining the columns on all pages of a dictionary, we proceed to the Slice stage. Within each column, we separate dictionary entries and their explanations from one another by means of digital boundary lines. In dictionaries such as Kâmûs-ı Türkî, where there is sufficient white space, we draw these boundaries as straight, unstepped, and level as possible.

In dictionaries such as the Lexicon (Redhouse) or Lugat-ı Remzî, where line spacing and entry layout are more compressed, we may need to use stepped boundaries. Even where stepping is necessary, we keep the boundary as close to a straight line as possible. We place the lines close to the text but never allow them to touch the characters. Under normal conditions we use horizontal and vertical steps, only in exceptionally crowded areas do we resort to diagonal lines.

7.1. Recording Continuity across Columns

At the Slice stage, we do more than separate entries from one another. We also identify textual continuity across columns. For example, a three-page set in a dictionary with two columns per page contains six columns in total. When moving from the first column to the second, we record whether an entry that begins at the end of the preceding column continues, or whether a new entry begins. We perform the same check at every column transition within the set. We also determine the relationship between the first column of the set and the last column of the preceding set, and between the final column of the set and the first column of the following set.

This information becomes crucial at the subsequent Highlight stage. If a dictionary entry extends across more than one column, we can combine all of its visual slices into a single group.

8. Third Stage: Highlight

Once Slice has been completed throughout the dictionary, we proceed to the Highlight stage. Here we isolate and categorize words and word groups within a dictionary entry that are significant for research purposes.

8.1. Headwords and Subheadwords

We first select the headwords. We also select as subheadwords structures that are generally formed by adding one or more words to a headword and that have their own independent definition within the dictionary entry. Subheadwords usually consist of a word group, although we also encounter single-word subheadwords. The essential criterion is that the expression in question has an independent definition of its own within the main entry.

Within an entry, we separately select words that provide important linguistic or lexicographical information about the headword. These may include expressions indicating the feminine or masculine form of the headword, its Chagatai equivalent, plural, singular, or synonym. These units are not subheadwords. They are nevertheless critical for understanding the dictionary entry and establishing relations between headwords, and we therefore highlight them under a separate category.

8.3. Examples

Words, word groups, or sentences that are neither subheadwords nor critical concepts but are supplied by the author as examples of the headword's usage are selected under the Example category.

8.4. Cross-references within the Dictionary

We also highlight, under a separate category, expressions by which the author directs the reader to another entry in the same dictionary. The notation used for such cross-references varies from dictionary to dictionary. In Kâmûs-ı Türkî, for example, the abbreviation با: may be used in the sense of “bakıla” (see). In Lehçe-i Osmanî, Ahmet Vefik Paşa may place the letter ن after the word to be consulted in the sense of “nazar oluna” (see). In the Lexicon (Redhouse), See is used, while Kâmûs-ı Fransevî may employ V. (Voir). Although their forms differ, their lexicographical function is the same, and we therefore treat them under the same functional category in the system.

8.5. The As Headword Category

Especially in Şemsettin Sami's dictionaries, we encounter structures that appear at first sight to contain more than one headword within a single entry. In some cases, different spellings or closely related forms with the same meaning are indeed grouped under a single definition. In other cases, however, the author does not present the forms listed side by side as independent headwords within that entry. Rather, he indicates that his preferred spelling appears elsewhere in the dictionary and directs the reader to that entry. Recording all of these forms as independent headwords within the same entry would therefore produce incorrect results. We consequently mark the first value as Headword and the others as As Headword.

As Headword indicates that the relevant value should in fact occur as an independent headword in another slice.

8.6. Highlighting Word Groups and Compound Words in Components

At the Highlight stage, we avoid selecting word groups and compound words as single blocks wherever possible. In the expression kelâm-ı câmi’ كلام جامع for example, we highlight kelâm كلام and câmi’ جامع separately and link them to one another. If the word group is the second subheadword in that entry, we may assign IDs such as 2a and 2b. In this way, we record in detail that they are the first and second components of the second subheadword.

We can apply the same method to a compound written as a single word. Even if karagöz قرەكوز is written as one word, for example, we may select its components separately and link them because this is necessary for analytical parsing.

If, however, the letters of the compound are joined in such a way that they cannot be separated visually, we select the whole form as a single unit and leave the analysis to the Type-O stage. Doing otherwise would distort the structure of the text and run counter to a methodological principle that we follow carefully throughout the project.

8.7. Highlighting Principle for Inflected or Affixed Compound Words

As noted above, when a compound word has not taken an affix, we select the constituent words separately and link them to one another. When an affix applies to the compound as a whole, however, we follow a different method. Consider the following example:

In this example, because the Arabic nisba suffix -i ی in karabâğî قرە باغی applies not merely to the component bâğ باغ but to the compound as a whole, we select the entire compound together with its suffix as a single unit at the Highlight stage. Otherwise, bâğî باغی would appear to be a component of the compound, which would introduce a serious error. There is, of course, an Arabic word bâğî باغی meaning ‘rebel’ or ‘one who rebels’, but it has no semantic connection with karabâğî قرە باغی.

Failure to follow this method can also generate meaningless forms. In the word, ihlâsperverâne اخلص‌پرورانە, for example, an incorrect segmentation would produce the meaningless word perverâne پرورانە. The Persian suffix -âne انە does not attach only to perver پرور, rather it attaches to the entire compound ihlâsperver اخلص‌پرور. To prevent spurious word formation and false semantic matches, we decided to highlight suffixed compound words as single wholes rather than dividing them, and we apply this rule consistently throughout the project.

8.8. Identifying Word Groups Not Explicitly Repeated in the Text

Historical dictionaries frequently write an element shared by several word groups only once. The author writes the common first word of several expressions that would ordinarily be written separately and then lists their second elements. In a conventional OCR approach, because the shared element that is not physically repeated in the text is recognized only where it actually appears, the other implicitly present word groups do not emerge in the digital version. As a result, those expressions cannot be retrieved in search. In BuildLQ, we explicitly record relationships that are clearly recoverable from context.

For example, if the text first gives the phrase nezâret-i celîle نظارت جلیلە and then mentions only the adjectives behiyye بهیە and aliyye علیە, and the context makes clear that the intended phrases are nezâret-i celîle نظارت جلیلە, nezâret-i behiyye نظارت بهیە, and nezâret-i aliyye نظارت علیە, we link the shared word of the phrase separately to each relevant adjective even if it is physically distant and other words intervene. The application preserves these multiple relationships through automatic ID assignment. The IDs maintain both word order and the links between components separated by one or more intervening words. In this way, for example, one explicit and two implicit word groups, or two explicit and several implicit ones, can be recorded as distinct data units and made searchable.

This method allows us both to preserve the source text without altering it and to represent digitally the lexical structure that the author did not explicitly repeat but that is present in context.

8.9. Creating Space for Words Described but Not Written

We occasionally encounter cases in which a word itself is not written, but its letters or manner of formation are described. If the described word is important to the headword, we draw an empty area at the relevant location during the Highlight stage. For example, in the entry ilhâk الحاق, after the second sense is explained as “to add to the end, to be added to the end” (ahirine getirmek, getirilmek) the author states that the letter ی may, when necessary, be added to دلجو. The second form itself is not written, the author merely describes how it is formed. In such a case, we create an empty area on the image for the described second form. We then enter the word in Ottoman script at the Type-O stage and in Latin script at the Type-L stage.

9. Fourth Stage: Type-O — Typing in Original Letters

Once all units selected during Highlight have been completed, we proceed to Type-O, or Typing in Original Letters. If we take a three-page set as an example, those three pages may yield six columns, approximately sixty dictionary entries, and, from within those entries, roughly two hundred selected units including headwords, subheadwords, critical concepts, terms, and word groups. At the Type-O stage, these words and word fragments selected at the Highlight stage become our units of work.

9.1. OCR and Human Verification

We first process the selected words using OCR tools and artificial intelligence. Our aim is to prevent people from having to retype from scratch values that the machine can already read correctly, thereby avoiding unnecessary loss of time. Beneath each selected word is a card with a wide input field in which the Ottoman-script value is entered. This field is automatically pre-populated with the OCR output, but the value is not saved automatically. The person working on the relevant set checks the OCR result. If it is correct, no action is required. If there is an error, it is corrected. If the word is inflected, plural, compound, or otherwise requires analysis, its morphological analysis is also carried out at this stage.

9.2. Separating the Author's Language Attribution from Our Language Analysis

One of our key principles at the Type-O stage is to keep separate the information supplied by the dictionary author about the language of a word and our own analysis of the language elements within that word. Some historical lexicographers may assign the entire word to a language on the basis of the suffix it takes. They may, for example, classify eczâcı اجزاجی as Turkish. We, by contrast, record all language elements within the word separately. Thus, if a form such as eczâcı اجزاجی contains both Arabic and Turkish elements, we activate both the Arabic and Turkish labels on our analysis side. If the author classifies the same word as Turkish, we preserve that separately by activating the relevant language label on the author’s side. We do not call the author’s classification ‘wrong’; we keep the author’s assessment distinct from the analytical classification adopted in our own methodology.

In this example, we see that the Persian word mûytâb مویتاب entered Turkish as mûtâf موتاف. As shown, the author of Kâmûs-ı Türkî records Persian as the word’s language of origin. To preserve this information, we activate Persian among the language labels on the right-hand side of the card. Since the word’s spelling underwent a transformation in Turkish, however, the new form also contains a Turkish language element. On the left-hand side, where we record our own view of the language elements, we therefore activate both Persian and Turkish.

9.3. Recording Misspellings

If we consider a word to be incorrectly spelled in the source image, we mark it as Misspelled. On the first card, we preserve the form actually present in the image and label it as misspelled. The system then creates a second card below it, where we enter the corrected form and continue the morphological and linguistic analysis on the basis of that corrected form. Our claim that the form is erroneous is restricted to the copy in our possession and, where necessary, even to the particular digital image we are using. We do not claim that the word is also incorrect in the author’s copy or in another edition. The error may have arisen in printing, in the copy itself, or through poor scanning or image processing during digitization. We record that the value appears in the image in the form shown on the first card and add below it the spelling that we consider correct. In this way, we preserve separately both the spelling found in the source and the form corrected through our own interpretation.

Identifying spelling errors is not always as straightforward as detecting a character that has clearly been written incorrectly. In Lugat-ı Nâcî, for example, under the headword karâr قرار, after the subheadword karâryâb قرارياب, we encounter a subheadword written as karar (‘decision’) but defined as “kararsız” (‘unstable, unsettled’). On the basis of this clear semantic contradiction, we infer that the subheadword should in fact be bî-karâr بی‌قرار. In such cases, we preserve karâr قرار, the form visible in the image, as the original value and record the corrected form bî-karâr بی‌قرار separately.

When the visitor sees the corrected form on the front end, hovering over the word opens a tooltip that also shows how the word appears in the source. If desired, the user can then click the corresponding original image to view the word first within the dictionary entry, then within the column, and finally within the full page. In this way, the user can move step by step from a corrected or analyzed value back to its original context in the source image.

For values marked as Misspelled, we also record the type of error. We use categories such as Letters for an incorrect letter, Harakat for incorrect vocalization, and Complicated where more than one element is problematic. In this way, we preserve in the database not only the fact that a spelling is erroneous, but also the level and type of the error.

9.4. Roots, Primary Value, and Inflected/Derived Structures

When analyzing a word, we remove its affixes step by step and record the resulting analytical values on separate cards. Among these cards, we designate as the Primary Value the form that is to serve as the principal lexical value of the word. The Primary Value card is displayed in the system with a green border.

When an independent word has been formed by derivation, we select the derived form itself as the Primary Value. In the first example, because hulûskârlık خلوصکارلق is a derived word, we designate the upper card outlined in green as the Primary Value and then descend to hulûskâr خلوصکار and ultimately to the root hulûs خلوص.

By contrast, if the upper form is merely inflected, we select the uninflected base form as the Primary Value. In the example below, erbaadan اربعه‌دن is an inflected form. Once the inflectional suffix is removed, the root erbaa اربعه becomes the Primary Value shown with a green border. In this way, the system records not only root-affix relationships but also which level in the analytical chain constitutes the principal lexical value.

9.5. Lexicalized Inflected Forms

Some words remain morphologically inflected but are nevertheless treated by the dictionary author as independent entries and assigned a separate meaning. We refer to such forms as Lexicalized Inflected Forms.1 Civârında جوارندە is a good example. Such forms differ from derived words. In kalemlik قلملك, for instance, the derivational suffix -lik لك creates a form distinct from kalem قلم, so it is unsurprising that the dictionary gives it an independent place and meaning. The interesting feature of civârında جوارندە is that the author has lexicalized a form that remains inflected. In such cases, we activate the Inflected label on the relevant card and, contrary to our general practice, designate the inflected form itself as the Primary Value. This allows us to display such values under Lexicalized Inflected Forms in the results list.

9.6. The Detailed Gradation We Use between Inflection and Derivation

I consider the conventional Turkish binary of inflectional suffix / derivational suffix somewhat too broad, especially on the derivational side, to capture important differences of level among suffixed words. I therefore prefer a more detailed classification in our analysis, dividing inflection into three levels.

The first level consists of forms that arise solely because of syntactic context. We place forms such as kalemi, kalemden, and kalemimi, produced by nominal case and related suffixes, at this level. We call this Contextual Inflection.

At the next level are inflections connected with inherent grammatical properties of the word, such as number, gender, and negation. We place relationships such as ağaç → ağaçlar, nebî → enbiyâ2, kerîm → kerîme, and çıkmak → çıkmamak in this group. We call this level Inherent Inflection.3

At the third level are inflections that change the word class but do not necessarily produce an independent lexical unit. Forms of a verb such as yap → yaparak, yapan, or oku → okuduğu are treated at this level. We call this Transpositional Inflection.4

Above these levels are derivational relations that produce a new lexical unit. We place examples such as süt → sütçü and ser → sergi in this group. The essential criterion is that the resulting form does not merely acquire a grammatical function, but carries a new and independent lexical meaning. The same form may therefore fall into different categories depending on its use. Yazma, for example, may be treated as Transpositional Inflection when used as a nominalized form of the verb yazmak; by contrast, lexicalized yazma in the sense of ‘manuscript’ belongs to Derivation because it has become an independent lexical unit. Likewise, we do not assign dondurma meaning ‘an act of freezing’ and dondurma used as the name of a food to the same morphological category.

9.7. Multi-layered Analyses

Some words require a multi-level analysis. In giyinti كییندی, for example, we first descend to giyin كیین and then to giy كیی. Because giyinti كییندی is formed through derivation, we mark it as the Primary Value. We also treat giyin كیین within a derivational relation. Since the final-level form giy كیی is a Turkish verb, we activate the Verb label.

With a word such as kehrubâiyet كهربائیت, the analysis can become a long chain involving the separation of several language elements and affixes. The initial form may contain Arabic, Persian, and Turkish elements together. When we move from the form with ت to the form with ة, the Turkish element disappears. We then remove the suffix that contributes the Arabic element to reach the Persian level, and subsequently remove the Persian relational suffix to arrive at kehrubâ كهربا.

We then divide the word into the components keh كه and rubâ ربا. Since there is also a lower-level value kâh كاه for keh كه, we record that relationship separately as well.

9.8. Fragments Split at Line End (Fragment)

A word may be broken at the end of a line, with one part written at the end of one line and the other at the beginning of the next. If the two parts together constitute a single word, we mark them as Fragment at the Highlight stage and assign them IDs such as 1a–1b to show that they are parts of the same word. At the Type-O stage, we combine these fragments to reconstruct the complete word.

9.9. Fragments of Constituent Words in Compound Structures (Compound Fragment)

Especially in dictionaries such as the Lexicon (Redhouse) and Kâmûs-ı Fransevî, we see authors use an em dash in a definition instead of repeating the headword. The mark may stand for the headword in its entirety. An affix and another word may then follow, producing a new word group. In such cases, we select the parts of the word group as Compound Fragments. A Compound Fragment may be classified as a complete word, Prefix, Suffix, Infix, or Meaningless. Because the em dash stands for the entire headword, we treat it as a complete word. For affixal fragments such as prefixes, suffixes, and infixes, we also record whether they are inflectional or derivational in nature.

We then combine these fragments to form the complete suffixed word, assign the language labels for the new level, and subject the word to our usual morphological analysis.

9.10. Analysis of Compound Words

In compound words, we first remove the word’s inflectional suffixes, then descend to the components of the compound and analyze each component separately. In birbirine بربرینە, for example, we first remove the inflectional suffix and arrive at birbiri بربری. We then divide this into the components bir بر and biri بری. After removing the suffix from the second component as well, two base values bir بر emerge. In this way, we can represent both components of the compound explicitly in the database.

In the example ihlâsperverâne اخلاص‌پرورانه, Arabic and Persian elements occur together at the first level. We descend from ihlâsperverâne اخلاص‌پرورانه to ihlâsperver اخلاص‌پرور, and then to the components ihlâs اخلاص and perver پرور, which we label as Root. By default, a lower card in the system is linked directly to the card immediately above it. If, however, a component needs to be linked not to the card immediately above but to a higher compound value, we select the relevant cards, manually change the relationship, and use the arrow to link the component to the correct upper level.

This allows us to list values such as ihlâsperver اخلاص‌پرور and ihlâsperverâne اخلاص‌پرورانه among compound structures when a user searches for ihlâs اخلاص.

9.11. Relationship of Lower Analytical Values to the Upper Value

The analytical value beneath a word is often the root of the word above it, but the relationship is not limited to this. The lower value may also be the singular, synonym, near synonym, or another related form of the upper word. We also use labels such as Uncertain Spelling, Feminine, and Masculine. For example, in the analysis of a word such as üflemek اوفلمك, if a letter not visible in the upper form reappears in the lower value because we consider it necessary in the root or imperative form — as when ه is added in üfle اوفلە — we label the lower form as Uncertain Spelling.

Similarly, in examples such as merhûm مرحوم → merhûme مرحومە or cemîle جمیل → cemîl جمیلە, we label the relevant words Feminine and Masculine, respectively.

9.12. Reducing Verbs to Their Base Form

For verbs recorded with the suffixes -me مە / -ma مە or -mek مك / -mak مق, we remove these suffixes to reach the base form of the verb and label that value Verb. We apply the same approach to other inflected verb forms. When analyzing forms such as geldiğini, gelmiş, gelmeyeceğini, or gelmeyebileceğini, we ultimately reach the root gel at the lowest level of analysis. Our aim is to bring the different inflected forms of the same verb together around the basic imperative form.

9.13. Our Language Policy for Words Borrowed into Turkish from Other Languages

At the Type-O level, we generally label words that entered Turkish from languages other than Arabic and Persian as Other. For words borrowed from Arabic or Persian, we examine whether a change in spelling has occurred. If a word such as kâdir قادر is written with the same basic spelling in Arabic and Ottoman Turkish, for example, we assign only the Arabic label. By contrast, in words such as devlet دولت → دولۃ, şevket شوكت → شوكۃ, ümmet امت → امۃ or millet ملت → ملۃ, the ة of the Arabic source form is represented by in Turkish spelling. Similarly, in words such as imzâ امضا → امضاء, duâ دعا → دعاء and ricâ رجا → رجاء the hamza present in the Arabic form may be dropped in Ottoman Turkish usage.

In such cases, we first record the form used in Turkish and mark both the Turkish and Arabic language elements. Beneath it, we separately enter the legitimate Arabic spelling, for example the form with ة or with ء.

If the change is purely phonetic, by contrast, we do not activate an additional Turkish language label at the Type-O stage. We handle sound differences such as the transition from Persian cüvân جوان to Turkish civan جوان at the Type-L stage.

9.14. The "Unknown" Language Label

We maintain Turkish, Persian, and Arabic as the system’s principal language categories. Because Kâmûs-ı Türkî contains frequent references to Chagatai, we have also defined Chagatai as a separate language category. Languages outside these categories are generally classified as Other. We also use the label Unknown. It does not mean ‘we do not know the language of this word’, rather, it records cases in which the language of origin is itself stated to be unknown. If this is the author's view, we mark it on the author side. If it is our own analysis, we mark it on the project side.

Here are two examples of words whose language of origin is explicitly recorded by the dictionary author as unknown. In the first, the author writes “its origin is unknown” (aslı meçhul); in the second, “its origin could not be determined” (aslı anlaşılamadı).

An example of a word whose language of origin we have been unable to determine:

9.15. Recording Sources Used in Analysis

When we draw on external sources in analyzing a word, we record those sources on the relevant card. We make particular use of Nişanyan Sözlük, Kubbealtı Lugatı, and reputable online or printed Ottoman dictionaries, especially the Lexicon (Redhouse) and Kâmûs-ı Türkî. Where necessary, we also add the source link, page number, and an explanatory note.

For example, if we move from besle بسلە, which remains in current Turkish usage, to the root besi بسی, we always ground this analysis in a source. Once the suffix -le لە is removed, the resulting form is bes بس, a word no longer in circulation or current use. The root besi بسی likewise requires further evidence, since its unsuffixed form bes بس may itself be the root. In cases of this kind, we make a point of avoiding unsupported conjecture. We either cite a reputable source or do not descend to that root.

9.16. Recording When Pronunciation Involves Interpretation

In Type-O cards, we also use a label called Interpretive Pronunciation (Pron?). If the pronunciation of a word is explicitly indicated in the source text, we do not activate this label. If, however, the word has no harakat or other explicit pronunciation marker and we read the value on the basis of our own knowledge of Ottoman Turkish, we activate the label even when we are highly confident in the reading. The label therefore indicates that the pronunciation involves some degree of interpretation.

On the other hand, when the harakat in the text indicate the pronunciation sufficiently, we do not adopt unnecessary skepticism. In Kâmûs-ı Türkî and Lugat-ı Nâcî, for example, dictionary authors explicitly explains how certain consonants that lack vowel markers are to be read and supplies harakat for the headwords. In such cases, we do not classify a reading based on the source's own notation as an interpretive pronunciation.

Nor do we claim absolute certainty about pronunciation. Even French- or English-language explanations written in Latin script do not always fully encode details such as open and closed vowels or intonation. Redhouse's use of four distinct phonetic characters for the sound ‘a’, whose precise phonetic values we still cannot determine with complete certainty today, is one example of this difficulty. The label therefore allows us to distinguish pronunciation explicitly encoded in the source from pronunciation that we have supplied through our own reading.

9.17. Recording the Author's Etymological Explanations

We also transfer into the data model etymological explanations supplied by the dictionary author within an entry. We make no claim, however, as to the correctness of these explanations. If, for example, Kâmûs-ı Türkî identifies aşı آشی as a noun and then gives the bracketed explanation aşmaktan آشمقدان, we link the word to the root aş آش and add the Verb label to the root. This enables us to distinguish the noun aş آش meaning ‘food’ from the verb aş آش meaning ‘to cross’. What we record here is the fact that Şemsettin Sami establishes such an etymological relationship in the relevant entry. LexiQamus does not thereby make an independent claim for the historical correctness of that etymology.

10. Establishing Lexical and Linguistic Relations at the Type-O Stage

We do not regard Type-O merely as a stage at which words are morphologically segmented. We also establish numerous linguistic relationships among headwords, their components, and the critical concepts selected within entries. A single entry may, for example, present more than one spelling of the same word as headwords. In the case of bileği بلگی, for instance, the other legitimate spelling of the word, بیلەگی, is written explicitly before the definition.

Within the entry, we may also select structures such as bileği çarhı بیلگی چرخی, bileği demiri بیلكی دمیری, bileği taşı بیلگی طاشی, or bileği kayışı بیلگی قایشی. From these expressions, for example, we analyze the form çarhı چرخی, remove the inflectional suffix, and reach the Persian root çarh چرخ.

Through the molecule icon that becomes active on cards selected as the Primary Value, we link that value to the relevant headword or headwords. We also establish linguistic links between headwords and the components of selected compound words and word groups within definitions. A point worth emphasizing is that we do not merely record that two concepts are ‘related’. We go further and specify the type of relationship. Even the direction of a relationship can change its type.

We use the following relationship categories:

  • Root

  • Singular

  • Same

  • Alternative Spelling

  • Near Synonym

  • Synonym

  • Feminine

  • Masculine

  • Inflected

  • Derived

  • Lexicalized Inflected

  • Plural

  • Closely Related

  • Moderately Related

  • Distantly Related

  • Antonym

  • Unrelated

  • Closer to Accurate

  • Less Accurate

  • Accurate

  • Corrupted

  • Misspelled

  • Counterpart

  • Common Root

  • More Used

  • Less Used

  • Mispronunciation

  • Misuse

  • Vulgar

We derived a substantial proportion of these categories from expressions used by dictionary authors themselves in their entry definitions. When we establish these relationships systematically across all dictionary entries, we will obtain connections among millions of words based not merely on superficial similarity but on detailed morphological and semantic categories. One of the major aims of this work is to bring words together into broad semantic and linguistic networks through these relationships.

11. Fifth Stage: Type-L — Latinization

After completing the analyses at the Type-O stage, we proceed to Type-L, or Latinization. On the Type-L screen, the Ottoman-script value entered in Type-O appears at the top. We treat each Ottoman-script value as a separate input unit. Thus, if a word has more than one root or analytical lower value, each appears separately at the Type-L stage. For every Ottoman-script value, we can enter one or, where necessary, more than one Latin-script value.

11.1. Shared Language Labels between Type-O and Type-L

We use the same language labels at the Type-L stage as at Type-O. These are not, however, two independent datasets. They represent the same language information on two different screens. If we notice during Type-L that a language label was assigned incorrectly at Type-O, we can correct it directly there, and the change is reflected in the other stage as well.

11.2. Spelling Pronunciation and Modern Turkish Pronunciation

When recording the Latin-script equivalent of an Ottoman-script word, we use two principal categories. The first is Spelling Pronunciation. We deliberately do not call this ‘original pronunciation’, because claiming that a word was in fact pronounced in precisely that way in a particular historical period or source language would go beyond what our evidence can support. By Spelling Pronunciation, we mean the pronunciation permitted by the existing spelling and represented, as far as possible, with the modern Turkish Latin alphabet. Our second principal category is Modern Turkish Pronunciation.

11.3. The Latin Character Set We Use

At the Type-L stage, we do not use a specialized and comprehensive academic transcription alphabet. The main reason is that we already display the original Ottoman-script value separately to the user. For Latinization, we rely primarily on the modern Turkish alphabet. In addition, we use ‘ for ع and ’ for ء, the straight apostrophe ' to represent ال constructions in Arabic compounds, and â, î, and û to indicate vowel length. This limited character set does not allow us to represent every phonetic detail completely, especially in Arabic and, to some extent, Persian words. As noted above, however, because we continuously display the original spelling as well, this does not create ambiguity.

11.4. Marking the Turkish Phonetic Element in Spelling Pronunciation

For example, even when we try to represent tama‘ طمع in modern Turkish letters as faithfully as possible to its spelling, we cannot fully represent the Arabic sound of ط in the modern Turkish alphabet. The sound ع likewise has no equivalent at all in Turkish. Since we use the dedicated character ‘ for ع, we can represent that sound separately in Latinization. We do not use a comparable special character for ط, however, and must therefore represent it with the closest Turkish equivalent. As a result, even when a Spelling Pronunciation is written as closely as possible to the original letters, elements of the Turkish sound system inevitably enter the Latinization. In such cases, we also activate the Turkish language label for the relevant Latin value. The same issue arises with other emphatic consonants in Arabic, interdental sounds, and other letters that have no exact counterpart in modern Turkish. The Turkish label here does not mean that the word is etymologically Turkish. It indicates that the Latin-script representation of the pronunciation contains a phonetic element belonging to modern Turkish.

In the example below, we activated the Turkish label because the Arabic phonetic values of و and ط in the word vatan وطن have no exact counterparts in modern Turkish.

11.5. Degrees of Probability in Modern Turkish Pronunciation

We can record modern Turkish pronunciation values at three levels of probability: Likely, Possible, and Unlikely. We have not yet applied this classification exhaustively to every word. At the present stage, we have primarily entered Likely pronunciations, while also adding Possible values in some cases. We plan to process lower-probability pronunciations more comprehensively in future work.

11.6. Near Match and Divergent

We also mark the degree of difference between the Ottoman-script value and the Modern Turkish Pronunciation. If the difference is limited to vowel-bearing letters or to changes between phonetically close consonants, we use the Near Match category. For example, in cüvân جوان → civan, tehlüke تهلكە → tehlike, and behâ بها → pahâ, the basic sound structure of the word is largely preserved and the change in pronunciation remains limited.

Where the sound structure changes more substantially, by contrast, we use the Divergent label. In the change çâryek چاریك → çeyrek, for example, modern Turkish pronunciation departs markedly from the pronunciation indicated by the Ottoman spelling, and we therefore record the value as Divergent.

11.7. Numerous Possible Pronunciations

We encounter considerable variation in the pronunciation of historical, religious, literary, or scholarly words that are not widely used in contemporary Turkish. Some words, such as تغلب, appear in written sources in numerous different Latin-script forms, for example, tegallüp, tagallüp, teğallüb, and tağallüb. A single word may thus be represented in eight, ten, sixteen, or even more different ways. One reason is the absence of a comprehensive inventory giving a standardized contemporary Turkish pronunciation for all historical Ottoman words. Differences in how vowel-bearing letters, ب, غ, and other letters are sounded can produce substantial variation in Latinization. In my observation, five factors are particularly important in these divergent pronunciations and spellings: (1) the geographical background and (2) the ideology of the person Latinizing the Ottoman word, (3) the source language of the word, (4) the period in which the Latinization was produced, and finally (5) the rules of Turkish vowel harmony. As noted above, at the present stage of the project we have primarily entered, and continue to enter, the most likely Turkish pronunciations. We plan to expand the less common pronunciations in future work together with probability labels.

11.8. Recording Pronunciations Indicated by the Author

If the dictionary author explicitly states how a word is pronounced, we record this information separately as the author's view. For example, the Persian word terâzi ترازو has the source-language form terâzû, while Hakkı Tevfik's Turkish-German dictionary explicitly states that the Turkish pronunciation is terâzi. On the basis of this information, we record terazi as the Modern Turkish Pronunciation and separately note that this pronunciation is confirmed by the author.5

11.9. Multiple Legitimate Pronunciations

In some entries, the dictionary author explicitly states that more than one pronunciation of a word is legitimate, but chooses to describe the additional legitimate pronunciations rather than mark them with harakat. In such cases, we enter these readings without activating the Interpretive Pronunciation label, because we are not inferring the pronunciation ourselves. Rather, we derive it directly from information provided by the dictionary author. In Kâmûs-ı Türkî, for example, Ş. Sami writes explicitly that the first letter of شجاع may be read with any of the three vowel marks and that two particular vowelings are possible in the plural.6

Even though the harakat themselves are not placed over the headword, we treat this explanation in the entry text as a direct source for pronunciation and create the corresponding Latin values accordingly. Moreover, because the author states that “…in the plural, it may be vocalized with either damma or kasra” (“…cem’inde damme ve kesre ile tahriki caizdir.”), we learn that شجعان may be read as şüc‘ân and şic‘ân, and we record both readings at the Type-L stage.7

In the example below, Muallim Naci gives the reading hasbe حصبە with harakat and then states in the definition that it “may also be read with a fatha on the sād” (“sad’ın fethiyle de lugattır”), indicating that hasabe is also a legitimate reading. We separately record the pronunciation that the author describes but does not write out explicitly.8

This makes it possible both (1) to access hasabe, the word's second legitimate reading, without first reading the definition and (2) to reach this specific entry in the dictionary when hasabe is searched in Latin script.

11.10. The Centrality of Ottoman-Script Values in the Search Architecture

In the LexiQamus data model, we establish all fundamental relationships through Ottoman-script values. Latin-script values are treated primarily as pronunciation representations of these Ottoman-script structures for display to the user. Thus, when a user searches in Latin script, the system first identifies the Ottoman-script value or values corresponding to the relevant Latin value and then conducts the search through those Ottoman-script values. We consider this approach more reliable and capable of producing a broader result set than establishing relations directly among Latin values. Otherwise, two words that are in fact related may fail to match because of different Latinization choices.

The reverse is also possible: two words that are not actually related may appear connected because of accidental similarity between their Latin values. We therefore use Latin-script search as an access layer to the Ottoman-script word network that forms the foundation of the data model. In other words, when visitors search for a word in Latin script, the system first identifies the corresponding Ottoman-script value or values and then conducts the search through those Ottoman-script forms.

12. Review and Quality Control Process

We do not treat data entry in BuildLQ merely as a production process, rather, we conduct it within a multi-layered review and quality-control mechanism. The system has three principal roles: User, Reviewer, and Admin. We do not, however, establish an absolute or permanent hierarchy, particularly between Users and Reviewers. A Reviewer's own sets may also be subjected to review when necessary. Especially when work begins on a new dictionary or a new stage of an existing dictionary, we can temporarily place the work of all users, including Admins, back under review.

12.1. Approval Required for New Users

When a user first joins the project, or when we have not yet seen enough evidence that their work is reliable, we place their sets under Approval Required. After the user completes and submits a first set, they cannot take another set until that submission has been approved. We continue this mechanism for a period of time. Once the user's submitted sets are being approved regularly and are no longer being returned for substantial corrections, we remove the Approval Required restriction. The user can then take new sets, complete and submit them, and continue working consecutively without waiting for approval of the previous set. Removing Approval Required, however, does not mean that we stop checking that user's work.

We continue to perform random checks at this stage. If we identify a significant problem in a set we inspect and return it to the user, the user is restricted again and cannot take a new set until the returned set has been approved. Once the user again begins to submit consistently error-free or nearly error-free sets, we can remove the restriction. This allows users whose reliability has been demonstrated to work without constantly waiting for approval while preserving the quality-control mechanism.

12.2. Determining Which Sets Are Subject to Review

A user's work can be kept in To Be Reviewed status without preventing that user from taking new sets. In this case, the user continues working without interruption while submitted sets enter the review pool. We can also enable review for a particular dictionary or working stage through the dictionary profile, causing the relevant sets to enter the review pool systematically. This distinction separates two different questions: (1) whether a user's previous set must be approved before the user can take another set, and (2) whether that user's work is to be examined separately by a Reviewer.

12.3. Reviewers' Working Pattern

Reviewers are not only users who perform their own data entry, they also review sets submitted by other users. To maintain a balance between data entry and review, we designed the system to require Reviewers to conduct reviews at regular intervals. After a Reviewer has submitted three sets, the system assigns a submitted set that requires review before providing another normal work set. In other words, we follow a principle of approximately one review for every three data-entry sets. The Reviewer receives the earliest suitable submitted set from among those produced by users whose work is subject to review.

12.4. Blind Review

The Reviewer does not see which user prepared the set under review. Our aim is therefore to ensure that the evaluation is based on the work itself rather than on the individual who produced it. After checking the set against the relevant rules, the Reviewer makes one of two principal recommendations: approval or return. The Reviewer's decision is not, however, the final decision. It is a recommendation that the set be accepted or returned.

12.5. Final Admin Review

Once the Reviewer has completed the review, the set moves to Admin control. The final decision is made by the Admin, who may agree with or override the Reviewer's assessment. For example, an Admin may approve a set that the Reviewer recommended returning, or return to the user a set that the Reviewer recommended approving. In this way, we use the Reviewer's judgment as an important layer of quality control while keeping the final decision centralized. Together with the final decision, we also write an explanatory note, particularly for problematic sets. Where we consider it useful, we record a screen video showing the errors and add a link to the video in the note. The system then automatically emails the user the approval or return decision together with this note.

12.6. Reviewer Status and Cross-review

In BuildLQ, we do not treat Reviewer and User as absolute quality classes or as fixed ranks. Work produced by a person with Reviewer status may also be subject to review and examined by other Reviewers. Because the same individuals can both enter data and conduct reviews, there is no permanent class of users whose sole role is to supervise others. Since the identity of the person who prepared a set is hidden during review, team members can, where necessary, blind-review one another's work.

12.7. Methodological Calibration for a New Dictionary or Stage

When we begin work on a new dictionary or move an existing dictionary into a new working stage, we may apply a more intensive review process to approximately the first 10–15 sets. During this period, team members review one another's work, including users whose reliability has already been demonstrated and who hold Reviewer status.

The purpose is not merely to identify individual errors, but also to determine whether different users make different decisions when confronted with the same dictionary structure or linguistic problem. When disagreement arises, we discuss the issue within the team, consider concrete examples, and reach a shared decision. We thereby do more than correct the set currently under review: we also establish the common policy to be applied in subsequent data entry.

For this reason, review in BuildLQ serves both quality control and methodological calibration. We use it not simply as a supervisory stage that identifies errors after the fact, but as a methodological mechanism that reduces interpretive differences within the team and helps us maintain a consistent analytical approach throughout a dictionary.

13. Conclusion

With the method we have developed in BuildLQ, we do not treat the digitization of historical dictionaries simply as the conversion of images into plain text through OCR. We distinguish the visual and structural organization of the dictionary page, the relationships between entries and columns, lexical units that are either explicitly written in the text or implicitly present in context, the components of words, spelling variations, language elements, morphological levels, and pronunciation values, and record each of these as separate data elements. One of the fundamental principles of this approach is to keep the source text itself separate from our linguistic analysis. We preserve the dictionary author’s view of a word’s language, origin, or pronunciation as a distinct piece of information and do not substitute our own methodological assessment for it.

Likewise, when correcting a spelling that we consider erroneous, we do not delete the form found in the source. We retain the original and corrected values separately. At the Highlight stage, we distinguish the headword, subheadwords, examples, critical concepts, references, and other lexical units that make up the dictionary entry, and define the relationships among them. At the Type-O stage, we analyze Ottoman-script values at the levels of word, root, affix, and component. At the Type-L stage, we relate these values to different pronunciation possibilities. In this way, we transform the dictionary into a data structure that is not only readable but also analytically queryable. The detailed linguistic relations we establish among headwords, subheadwords, compound words, phrases, and critical concepts within entries bring dictionary entries that appear independent into a broader lexical network.

By systematically recording relations of root, singular, plural, alternative spelling, synonymy, near synonymy, inflection, derivation, gender, common root, and the other relation categories, we can connect millions of words not merely through the forms in which they are written, but through detailed morphological and semantic relationships.

Our review mechanism is also an important part of this structure. We closely monitor the work of new users, continue to check the work of experienced users randomly or systematically, subject Reviewer assessments to Admin control, and, especially when beginning a new dictionary or working stage, turn differing interpretations within the team into shared policies.

Quality control is therefore not treated merely as an inspection carried out at the end of data entry, but as a continuous process that also makes the methodology itself more consistent. Because the relationships established across the system are structured data rather than free-text notes, they can later be used directly in search, filtering, grouping, and analytical operations.

I consider one of the innovations of this work to be its ability to connect, within a single continuous data structure, every stage from the visual source of a dictionary down to its smallest analytical linguistic units. To the best of our knowledge, there is no other project of comparable scope that progressively decomposes a dictionary through interconnected stages from page to column, column to slice, slice to selected lexical units, then to Ottoman-script values and finally to their Latin-script counterparts, while preserving the links among all of these stages without interruption. An Ottoman-script or Latin-script word is not stored merely as an independent textual datum: the database also keeps it traceable to the precise region of the source image, the slice to which that region belongs, and ultimately the page and volume from which it derives. At the same time, we analyze Ottoman-script words, where necessary, down to their roots, affixes, and components, recording the morphological, semantic, orthographic, and etymological relationships among them. At the Type-L stage, we represent different readings and pronunciation possibilities within detailed linguistic categories.

This approach makes contributions across several interrelated fields. At the broadest level, BuildLQ offers a model within the digital humanities for the detailed and relational digitization of historical sources. A more specific area of application is digital lexicography: rather than merely converting historical dictionaries into searchable texts, the application also transforms their internal structural and linguistic relationships into machine-processable data. More specifically still, the method offers new possibilities for digital Ottoman studies and the digitization of Ottoman dictionaries. However, I believe that the model is not limited to Ottoman Turkish. Applied to historical dictionaries in other languages, it could similarly make visible lexical, morphological, and semantic connections among the millions of words they contain.

The method is not limited to dictionaries. Its principal innovation lies in a modular and extensible data architecture that begins with the source image, progressively divides the material into smaller and more meaningful units, and preserves the connections among them throughout the process. Because this structure is not tied to any particular language or type of text, it can be adapted to historical texts belonging to a wide range of linguistic and writing traditions, including Ottoman Turkish, Arabic, English, Italian, Chinese, and others. No matter how large the database becomes, the ability to move easily up and down among page, column, slice, word, component, and character levels allows highly detailed data on a very large scale to remain manageable. New relationship types and analytical categories that were not anticipated in the system's initial design can also be added later without requiring changes to the existing database structure. In this respect, the data architecture provides a flexible foundation for developing new layers of analysis.

Alongside this architecture, the second principal foundation of the project is its multi-layered review and quality-control system. Data entered by Users are examined by Reviewers and, where necessary, returned with explanations or recommended for approval, the final decision is made by an Admin. At the same time, we do not establish an absolute class hierarchy among User, Reviewer, and Admin roles that would exempt the data produced by any one of them from scrutiny. Data produced by Reviewers, and even by Admins, can themselves be reviewed by other members of the team. Because reviews are blind, the Reviewer does not know who prepared the set and can therefore assess the work of another Reviewer at the same level, an Admin, or a User according to the same criteria. By systematically managing this cross-review process, the application makes data quality less dependent on individual expertise or personal attentiveness and turns it into a methodological safeguard.

I believe that the combination of this detailed, traceable, and extensible data architecture with a multi-layered quality-control mechanism represents a significant methodological advance specifically for the digitization of historical dictionaries. More broadly, the same approach can contribute to digital humanities, dictionary studies, and lexicography by enabling historical dictionaries and other historical texts in the world's languages to be transformed into richer, queryable, and interconnected digital resources.

Dr. Ahmet Abdullah Saçmalı

Üsküdar, 15 September 2026

Acknowledgements

LexiQamus is a ten-year story. The project will probably continue for at least another forty years and will come to an end only when we finally reach the shores of the vast sea that is the Turkish language. That figure is not meant as hyperbole. If I live long enough, I hope to uncover all the words that remain hidden in the written sources of the language and to produce a work that encompasses its entire vocabulary.

This journey began in 2015. We released the first version of LexiQamus in 2016 and the second in 2020. Now, with LexiQamus 3.0, we are publishing its third major version.

Of course, I was never alone in this endeavor. Many friends, colleagues, and teachers have supported the project and contributed to its development.

First, I would like to mention my brother, Assistant Professor Muhammet Habib Saçmalı of the Division of Early Modern History in the Department of History at Marmara University. In addition to devoting an extraordinary amount of time and effort to the project, his critical interventions and insights prompted major changes in some of our fundamental policies. Beyond his training as a historian, it is difficult to overstate the value that his command of Arabic, Persian, English, and German has brought to a linguistic project of this kind. His sharp intellect, combined with an exceptional ability to immerse himself completely in a problem, played a vital role in formulating our data-entry rules, developing BuildLQ, and shaping LexiQamus 3.0. More recently, his contributions to digital design have been remarkable. His influence can be seen throughout both the front end and the back end of LexiQamus.

When we first began working with BuildLQ, people from many cities and universities across Türkiye joined the project. Among those who reviewed the early work, provided feedback to contributors, and helped shape our data-entry policies were Mehlika Çakmak and Furkan Arslan from the Department of History at Boğaziçi University, and Şeyma Kasapoğlu from the Department of Turkish Language and Literature at the same university. Despite their young age, they did work of considerable importance. During this formative period, we collectively made decisions about how the system should operate and laid the foundations for many of the data-entry policies that we still follow today.

The project later continued its journey with a new team.

Buse Büyükkeskin graduated at the top of her class from the Department of Turkish Language and Literature at Boğaziçi University and recently completed her master's degree in the same department. At our first meeting, because she preferred to remain relatively quiet, I must admit that I did not yet have a clear sense of her. Toward the end, however, she said, “I have a few notes,” and began sharing the points that had caught her attention. Not one of them was trivial. Each led us to make meaningful improvements, whether in the software or in our data-entry practices. In the years that followed, her suggestions on the Turkish language, software, and the overall direction of the project became central to the shaping of LQ 3.0.

Hatice Tüfekci, who received her bachelor's degree from the Department of Turkish Language and Literature at Sivas Cumhuriyet University and completed her master's degree in Old Turkish Literature at the same institution, has been with us since the digitization of Redhouse. She is the longest-serving member of the team. Her strong background in Turkish literature, extensive experience in textual reading, and knowledge of Persian have enabled her to make highly valuable contributions to the scholarly side of the project. She has also always been exceptionally attentive in identifying software bugs and diagnosing parts of the system that were not working as intended. Her willingness to spend long hours on even the most painstaking tasks is perhaps the clearest indication of her persistence and dedication. Equally valuable has been her ability to maintain that level of concentration over extended periods, notice subtle details that are easy to miss, and pursue them with great care.

Seher Bulut Köse, a graduate of the Faculty of Theology at Uludağ University who has also worked in the field of Turkish Islamic Literature, is an exceptionally meticulous and industrious member of our team. Her ability to manage several kinds of work successfully at the same time has allowed her to contribute to the project in many different ways. Having grown up with both Arabic and Turkish and using both as native languages, she has become a cornerstone of our work and an indispensable member of the project.

For example, when we encounter a phrase in Lugat-ı Ebuzziyâ that appears seriously problematic from the standpoint of Arabic, we are not always in a position to make a definitive judgment ourselves. Seher's contribution has therefore been particularly important in difficult questions involving Arabic. When establishing our policies on such matters as hamza, tāʾ marbūṭa, and yāʾ with hamza, Seher has consistently been one of our principal points of reference. The fact that she has been able to contribute to the project with such effectiveness and productivity while also caring for her children and family deserves special admiration.

Zeyneb Odabaş completed her undergraduate education at the Faculty of Theology at Marmara University and recently received her master's degree from the Department of Qur'anic Exegesis at the Social Sciences University of Ankara. Throughout our work together, she repeatedly demonstrated that she is both an excellent team member and a highly capable researcher who is prepared, when necessary, to defend with confidence what she believes to be correct. Her strong training in Islamic studies proved extremely valuable to us. During the Latinization stage, for instance, her subject knowledge often guided us in determining the possible pronunciations of particular Arabic letters. Alongside this scholarly expertise, she also worked with great precision on technical tasks such as image processing, dividing pages into columns, and slicing those columns. She brought the same care to the sets she reviewed. In short, she became one of the people whose judgment and work I trusted most on both the linguistic and technical sides of the project.

I am also deeply grateful to Dr. Enis Tombul, another graduate of the Department of Turkish Language and Literature at Boğaziçi University. Every question he raised and every objection he made was valuable. His expertise in the field and substantial experience in close textual reading contributed greatly to our work. Time and again, we saw that something that initially struck us as unusual might not seem unusual at all to someone who worked directly with such texts. In this respect, Enis's experience and judgment were especially valuable to us.

As the list below shows, the person who has so far made the greatest contribution to BuildLQ is Dr. Serap Arslan, who earned her doctorate from the Department of Turkish Language and Literature at Boğaziçi University. Her deep command of the field, her willingness to engage with difficult and highly detailed questions, and her role in resolving complex problems have been immensely valuable.

I also owe a special debt of gratitude to my dear friend John Zacharias Crist, who had already supported us during our work on Redhouse. Zack is an accomplished linguist and a true polyglot. Although English is his native language, he would engage with us in detailed discussions of Ottoman Turkish and offer remarkably perceptive observations on the finer points of Arabic and Persian. His advanced linguistic abilities have made a substantial contribution to the project.

I would also like to thank Fatma Nur Önür of the Department of Turkish Language and Literature at Istanbul University and Sümeyye Yıldırım, who holds a BA in Persian Language and Literature and is pursuing doctoral research in Turkish-Islamic Literature at Istanbul University, both of whom made important contributions over the course of the project.

My sincere thanks also go to Kevser Akın, a graduate of the Department of Turkish Language and Literature at Sakarya University, for her valuable work, particularly in the Split and Slice stages, and for patiently putting up with my demands throughout the process.

I would also like to thank Eren Aras Aydın, an engineer with a keen interest in complex linguistic questions, who distinguished himself during the project especially through the remarkable speed of his work.

Finally, I owe a debt of gratitude to all my students, younger colleagues, friends, and teachers whose names I have not been able to mention individually above, but who contributed to the project by offering advice, sharing their views when I consulted them, or even, sometimes unknowingly, introducing an idea or concept in the course of a brief conversation at a conference.

Contributors to LexiQamus 3.0

The following individuals are listed in descending order of contribution, taking into account the scope, nature, and intensity of their work on dictionary digitization and data entry for LexiQamus 3.0:

  1. Serap Arslan

  2. Zeyneb Odabaş

  3. Buse Büyükkeskin

  4. Seher Bulut Köse

  5. Hatice Tüfekci

  6. Sümeyye Yıldırım

  7. Habib Saçmalı

  8. Fatma Nur Önür

  9. Eren Aras Aydın

  10. Mehlika Çakmak

  11. Enis Tombul

  12. Zeynep Kılıç

  13. Zeyneb Bulut Erkuş

  14. John Zacharias Crist

  15. Kevser Akın

  16. Derya Doğan

  17. Munise Saydemir

  18. Eda Tuncer Yiğittepe

  19. Fatih Safa Memiş

  20. Şeyma Kasapoğlu

  21. Dilan Adanç

  22. Şehnaz İyibaş

  23. Bahattin Karakaya

  24. Senanur Cimitoğlu

  25. Furkan Akyıldız

  26. Hacer Er

  27. Merve Beyinli

  28. Esma Arzu Yetimova

  29. Fatma Altan Yılmaz

  30. Meryem Şentürk Çoban

  31. Hümeyra Yemenoğlu

  32. Adem Çakmak

  33. Mustafa Küpçü

  34. Alper Balcıoğlu

  35. Gülistan Ersöz

  36. Mehmet Akif Yılmaz

  37. Sevim Alkan

  38. Hasan Torun

  39. Gizem Pınar Civan

  40. Fatma Ersöz

References

Booij, Geert. “Inherent versus Contextual Inflection and the Split Morphology Hypothesis.” In Yearbook of Morphology 1995, edited by Geert Booij and Jaap van Marle. Dordrecht: Kluwer Academic Publishers, 1996.

Galancızade Hakkı Tevfik. Türkçeden Almancaya Lügat Kitabı. Istanbul: Matbaa-i Amire, 1907.

Haspelmath, Martin. “Word-class-changing Inflection and Morphological Theory.” In Yearbook of Morphology 1995, edited by Geert Booij and Jaap van Marle. Dordrecht: Kluwer Academic Publishers, 1996.

Ittzés, Nóra. A magyar nyelv nagyszótárának lexikográfiai koncepciója, különös tekintettel a szemantika és a grammatika összefüggésére a szótárírásban [The Lexicographical Concept of the Comprehensive Dictionary of Hungarian, with Special Reference to the Interrelatedness of Semantics and Grammar in Lexicography]. PhD diss., University of Szeged, 2011.

Muallim Naci. Lugat-ı Nâcî. Vol. 1. Istanbul: Matbaa-i Amire, 1901.

Oja, Vilja, and Iris Metsmägi. “From a Dialect Dictionary to an Etymological One.” Proceedings of the XVI EURALEX International Congress (2014).

Şemseddin Sami. Kâmûs-ı Türkî. Vol. 1. Istanbul: İkdam Matbaası, 1899.

Back to top