MEJE PROCESS · MEJE Librarying Workflow (21 chapters)
Chapter 5. Grouping into Wiki Headwords — Decisions in Second Integration
Chapter 5. Grouping into Wiki Headwords — Decisions in Second Integration
When first extraction ends, one table of 5,000 to 15,000 rows remains. It has many lines and nine categories, but it cannot be placed directly into a Vault. The same character is scattered across three rows under three different names; the same place appears in two forms; and the connections between keywords have not yet been organized.
Second integration is the work of arranging these scattered rows into one dictionary.
Among Librarying’s four stages, second integration is the stage most deeply touched by human hands. Its acceptance threshold rises from 60 percent in the first stage to 95 percent.
What Is a Headword? The Foundational Concept of Lexicography
To understand second integration, we must first grasp precisely the concept of a headword (headword, lemma).
Open an ordinary Korean dictionary and forms such as “went,” “having gone,” “going,” and “will go” do not appear as separate entries beneath the headword “to go.” They are all handled under that one entry. A headword is the standard form representing every form of a lexical item, and this principle is what makes a dictionary a dictionary. If inflected forms were separate entries, a dictionary would become dozens of times larger and a reader would first have to know which form to search in order to find information. The headword gathers every form in one place.
This is exactly what Librarying’s second integration does. If one character in “Seongra Gangho” is scattered in the first stage across four forms—given name, Chinese-character form, office title, and factional appellation—the second stage brings the four together under one headword. Whatever form is searched or encountered, all information becomes available at one headword.
Lexicography has principles for selecting a representative form: the most frequent form, the most basic form, or the form set by normative orthography becomes the headword. IP Librarying follows the same principle, making the form most frequently occurring in the IP or formally set by its producer the representative keyword. This is not an arbitrary decision about which notation to use. All future IP references will be made against this headword; once it is selected, the entire Vault is structured around it.
An entry from Dungeons & Dragons’ Forgotten Realms makes the principle clear. The character Elminster’s given name is Elminster Aumar; his appellations include the Sage of Shadowdale, the Old Mage, and several disguise names. One Elminster page in the Forgotten Realms Wiki gathers every form in one place, makes the given name representative, and leads a search for any appellation to that page. This is the structure second integration creates at the scale of 1,000 to 2,000 headwords.
Headwords and Knowledge Organization Systems
The sense of separating headwords from aliases has also long been addressed in library science and knowledge organization systems. W3C’s SKOS is a model for expressing controlled vocabularies such as thesauri, classification schedules, and subject-heading lists. It gives one concept a preferred label, keeps alternative labels and hidden search terms separately, and distinguishes broader, narrower, and related concepts.
Librarying’s thirteen columns share this sensibility. The representative keyword is the preferred label for a headword; alternative names and aliases are alternative labels that let users find the same object; related keywords are connections through which headwords depend on one another in meaning. Where a broader and a narrower category are clear, the relation can be read as hierarchical; where neither is above the other, it is better left as a related relation.
Librarying is not, however, an implementation of SKOS. It does not publish RDF or mechanically map to external knowledge systems. It uses the same principles at a lower level inside a creative IP’s Vault. The comparison is useful because it explains why aliases are kept separately, why related keywords must not be added carelessly, and why identical forms must be separated when they refer to two concepts. A headword is not a word list; it is the first unit of controlled vocabulary.
The Structure of First-Stage Results
Open the first-stage result table again. In the character category of “Seongra Gangho,” one swordsman is scattered across four rows: given name (“Seo Ungyeong”), Chinese-character form (“徐雲卿”), factional appellation (“Cheongeomsu swordsman”), and a judgment from the martial world (“the swordsman who finds an auspicious day hidden within an inauspicious day”). This condition—one person appearing by personal name, Chinese characters, affiliation, and sobriquet—is no different in Korean IPs and foreign works.
The same pattern appears in other categories. In mechanisms, “celestial energy,” “energy borne on starlight,” and “energy descending from the sky” appear separately. In places, an office’s formal name, abbreviation, and in-story designation appear separately. In objects, an organization’s formal name, acronym, and colloquial name appear separately. This scattering is the intended result of first extraction: it drew out forms as written and postponed deciding whether they meant the same thing. This chapter makes that decision.
Second integration begins by grouping rows with the same meaning. Group the four Seo Ungyeong rows under one headword, and they become one row: set “Seo Ungyeong” as the representative keyword; gather “徐雲卿,” “Cheongeomsu swordsman,” and “the swordsman who finds an auspicious day hidden within an inauspicious day” in the alias column; and combine their brief explanations into one-line definition. This grouping is the essence of second integration. It reduces 5,000 to 15,000 rows to 1,000 to 2,000 headwords, and each headword comes to contain every form with the same meaning. Only then do an IP’s materials begin to take the shape of a dictionary.
Works in which one being has several simultaneous forms show why this grouping matters. In Omniscient Reader’s Viewpoint, a constellation is called only by a descriptive epithet instead of its true name, hiding its identity; decoding that epithet reveals an actual mythological figure. The epithet “Demon-like Judge of Fire,” for example, is the archangel Uriel. Since one being has an epithet, true name, and mythological source at once, those forms must be grouped under one headword or readers will mistake the same figure for different people.
Why the Acceptance Threshold Is 95 Percent
The first stage accepts 60 percent, while the second leaps to 95 percent. The source of that difference matters.
A row misclassified in first extraction is acceptable because it passes through second integration once more. Whether it entered characters or objects, a person sees it again in the second stage and corrects the category. A row grouped incorrectly in second integration, however, becomes an .md file unchanged during third skeleton building. If two characters are grouped under one headword, their information is mixed into one .md file. The error then enters fourth-stage narrative writing and remains until verification finds it; the cost of recovery after one wrong grouping is high.
This is why the second-stage threshold is 95 percent. It is not 100 percent because human decisions also occasionally miss. That remaining five percent is caught in final verification, making 95 percent the most demanding line people can reach. This rigor makes second integration human-led: an AI work partner may rapidly generate synonym candidates, but people make the last decision.
The Structure of a Thirteen-Column Headword
Where the first-stage output had five columns—IDX, category, keyword, description, and source—the second-stage output has thirteen. Its foundation information from the first stage—IDX, category, and source list—survives, and eight columns are added.
Columns that establish a headword’s identity. The representative keyword is the form that represents the grouped synonyms. Use the IP’s official name or most frequently occurring form. This becomes the headword file name, and other Vault headwords cite it as [[representative keyword]].
Once the representative keyword is set, all Vault references use it. If an IP production team formally changes a device name, changing the representative keyword to the new name and retaining the old one as an alias automatically updates all wikilinks in the Vault. This is the strength of a single source of truth.
Worldview axis is a second classification criterion orthogonal to the nine categories. It answers, “To what structure of our worldview does this keyword belong?” Chapter 6 treats it in detail. Alternative names and aliases are forms that refer to the same object besides the representative keyword. They are a search column: they must be registered so a reader who searches an IP wiki by alias reaches the headword. A sister workflow examining a new short-story manuscript also needs aliases registered in order to recognize automatically which headword a form belongs to.
Multilingual columns. The English name is a proposed English notation for translation, and romanization follows Revised Romanization (RR) for Korean. Names use forms such as “Min-ji”; place names use romanization plus type, such as “Haewon Restaurant”; culturally specific terms use transliteration plus explanation, such as “Han (lingering resentment).” These two columns are core resources for IPs preparing multilingual publication. They are the first columns translators consult on opening a Vault, eliminating unnecessary time spent independently deciding spellings.
Body columns. A definition is the most compressed, self-contained one- or two-sentence statement of what the headword is. It is written without [[links]].
Detailed description is body text of two to eight sentences containing operating principles, history, structure, and relations to other headwords. It eventually contains [[links]], but a worker does not insert links manually from the beginning. Write the detailed description first as pure text; in the second-stage merge, related headwords are automatically wikilinked using the representative-keyword and alias lists. This prevents arbitrary alias links such as [[X|Y]] from mixing in and keeps the entire Vault ordered around representative keywords. Draft definitions and detail need only reach about 80 percent in the second stage. Unlike a core decision such as grouping synonyms, which must exceed 95 percent, definition and detail are refined again during fourth-stage narrative writing, so correct direction and core content are enough.
Related keywords lists other headwords connected to this one. Once filled, third skeleton building automatically injects wikilinks, drawing links between nodes in Obsidian. This column builds the keyword cloud’s network; a headword with no related keywords remains an isolated node in graph view. The denser this network, the better the Vault.
Period and translation metadata. Version status is the temporal status of a headword. “Current” means a headword now in use in the IP and applies to most entries. “Past” means it was used at one point but its notation or meaning has since changed. A place name used in an IP’s first edition and changed in its third is an example.
“Retired” means no longer used. If the reason for retirement is recorded in the translation note, the next person does not create confusion by using the same form again. “Dual notation” is for two forms intentionally coexisting, as when regions call the same thing differently. “Undecided” is for a headword not yet confirmed. Past and retired headwords are preserved rather than deleted because they appear when older works are read; a reader who sees an entry in the first edition and searches it must be able to find the information.
Translation note records cultural and contextual cautions for translation. For culturally specific terms such as han (恨), record transliteration and explanation; for shamanic terms, use forms such as “Mudang (Korean shaman)”; for period vocabulary, record the context that translation must heed. It is a letter to future translators, so the same decision need not be made again in the next translation job.
These thirteen columns make one headword row. Once 1,000 to 2,000 headwords are ordered, the whole vocabulary of an IP fits in one table.
What People Decide
Let us identify what people decide directly within the thirteen columns.
The most difficult decision is synonym grouping. If Seo Ungyeong in “Seongra Gangho” appears as a given name, Chinese-character form, factional appellation, and martial-world sobriquet, someone must decide whether all four belong under one headword or whether one form refers to another person. The decision appears obvious to someone familiar with the IP, but the same factional form can refer to different people. “Cheongeomsu swordsman” applies to both Seo Ungyeong and Samok: one is an orthodox swordsman who draws a sword on an auspicious day governed by auspicious stars; the other receives the texture of inauspicious stars as well and enters an unorthodox path. If two people who share Cheongeomsu as a naming quality are grouped under one headword, an orthodox swordsman’s definition mixes with an unorthodox swordsman’s. A person must decide in which of the two headwords the one form “Cheongeomsu swordsman” belongs.
Notations can also change meaning over time. A device name may indicate one thing in the first edition and another in the third. Someone must decide whether the two historical uses belong under one headword or two. This book recommends splitting them and marking version status “first edition only / third edition only.” Keeping them together makes the two meanings collide in one file.
If a character with the same name is a minor role in season one and a protagonist in season three, it is more appropriate to keep one headword and organize the chronological changes in the body. Such judgments can be made accurately only in the mind of someone who has read five years of work.
Headword naming is the decision, after forms of the same meaning are grouped, about which form becomes representative. In a Korean IP, choose the official Korean name; in a Korean publication of a foreign work, choose the notation that publication will use; where no official notation exists, choose the most frequent form. Once the representative is chosen, every other form becomes an alias. Aliases are searchable but do not appear as the headword title, and Vault wikilinks converge on the representative keyword.
Worldview-axis decisions and version-status decisions also belong to people. The permitted worldview-axis values—five to seven axes per IP—are fixed in the profiling stage before first extraction. AI can propose candidates, but judging which axis fits an IP’s structure belongs to people. Only someone who knows the IP’s time flow can decide version status. The fact that a form introduced in edition one was retired in edition three lives in the mind of someone who has read both editions; even if AI has read both, it may not distinguish intentional retirement from a mere notation change.
Definition acceptance is also a human judgment. People assess whether a one-line definition assembled from several short explanations passes: too short and the headword has no outline; too long and it becomes extended narrative rather than definition. One to three lines is appropriate.
Delegating to AI Work Partners
AI work partners work quickly where they support human decisions. Synonym-candidate extraction takes the first stage’s 5,000 to 15,000 rows and creates candidate groups of rows likely to mean the same thing, grouping forms that look alike, occur in the same source, or share a definition. Roughly 70 to 80 percent of these candidates pass as they are; people refine only 20 to 30 percent. A one-line definition draft is made after an alias group is settled: it combines short explanations collected in the first stage and compresses them to one or two lines, which people review to about an 80 percent level. AI can propose related keywords, but final wikilink injection is more stable when automatically processed against the global headword list. For worldview-axis candidates, give the permitted values fixed in profiling and attach the closest-axis candidate to each headword for human review and confirmation.
In terms of input and output, the input is one first-stage five-column CSV, the IP’s synonym-grouping guide, its permitted worldview-axis list, and the format of a thirteen-column integrated CSV. The output is a thirteen-column integrated CSV with synonym-group candidates attached, one-line definition drafts filled, related-keyword wikilinks automatically injected, and worldview-axis candidates attached. The intent of delegation is both speed and consistency. It would take a week for a person to read 5,000 rows once and construct synonym candidates; candidates arrive in about thirty minutes, reducing human decision time to two or three hours. The expected result is a thirteen-column CSV of 1,000 to 2,000 ordered headwords in roughly two to four hours. The limits are the final synonym decision, deciding whether to create a worldview axis, the tone of definitions, and version status; people must decide all four.
The Flow of Decisions
The second-integration workflow is this. First, deliver the entire first-stage five-column CSV to an AI work partner together with the synonym-grouping guide, thirteen-column format, and the list of permitted worldview-axis values determined in profiling. For roughly thirty minutes to an hour, AI creates synonym candidate groups, one-line definition drafts, related-keyword candidates, and worldview-axis candidates. In the merge stage, headwords in detailed descriptions automatically become [[representative keyword]] links using the representative-keyword and alias lists, and values outside the permitted worldview axes are caught as errors.
The person receiving the result begins review against the demanding 95 percent threshold. They inspect synonym groups one by one to decide whether they are correct, should be split, or need another row; review and confirm worldview-axis candidates; fill version status; and review definition drafts.
Review takes roughly two to three hours, enough to pass through an IP’s 1,500 headwords once. Depth differs by headword: 100 to 150 hub headwords receive one to two minutes of decision each, while 1,000 leaf headwords take about ten seconds each. When automated counts are accurate, leaf decisions mostly only need approval.
When review ends, the thirteen-column integrated CSV has crossed its acceptance threshold. The first stage’s 5,000 to 15,000 rows have been grouped and ordered as 1,000 to 2,000 headwords, and the IP’s materials have finally reached the shape of one dictionary.
What Becomes Visible After Second Integration
Second integration reveals an interesting result. Sort the thirteen-column CSV by each headword’s citation count: the automatically counted number of times a headword occurs in other headwords’ related-keyword columns. Put the most cited headwords first and inspect the top ten, and what the IP places at its center becomes visible at a glance.
In “Seongra Gangho,” that ordering placed “celestial energy,” the core resource determining martial power, first; “one who defies heaven,” the protagonist and sole exception to the celestial-energy order, second; “Martial Alliance,” the center of power controlling the calendar order, third; and “orthodox and unorthodox factions,” which divide the martial world, fourth. The IP’s identity is compressed into these four headwords. After five years of operation, it became apparent that its identity had never before been confirmed so compactly. When 1,500 headwords are organized, an IP’s identity appears in its hub ranking; this is one expected benefit of Librarying.
There is a trap here: grouping synonyms too broadly. A clear foreign-work example comes from Pride and Prejudice. The Bennet family has five sisters, and the convention of address allows any of them to be called “Miss Bennet”; an occurrence of “Miss Bennet” does not automatically mean the eldest daughter, Jane. The recommended approach is to hold the title separately when first encountered, group it as an alias only where the text makes the person explicit, and retain unmarked occurrences as “Miss Bennet (unspecified)” for another decision in fourth-stage narrative writing or verification. A wrong grouping allows one character’s definition to invade another’s.
The same trap appears where a single title points to different people over time. In Solo Leveling, “Shadow Monarch” is an office title; the first Shadow Monarch’s proper name is Ashborn, and protagonist Sung Jinwoo inherits the power and becomes the second Shadow Monarch. The title therefore points to Ashborn and Sung Jinwoo according to the work’s moment. The headwords must record the title, proper names, and succession relation together so it is possible to tell whom “Shadow Monarch” means in each passage.
Automated counts have limits too. They count occurrences in related-keyword columns and mark high counts as hub candidates, but frequency is not identical to an IP’s core. In “Seongra Gangho,” “calendar law,” a criterion used to measure the strength of nearly all martial artists, has a very high citation count yet is closer to background material than the IP’s essence. Conversely, a headword with few citations may determine the IP’s identity. “Nameless star,” belonging to no constellation, is such a case: it appears rarely, but that absence itself makes up the IP’s essence. When automated counting marks hub candidates, people confirm which ones are actually central to the IP. Chapter 8 examines this judgment in detail.
Closing Second Integration
Scan a thirteen-column CSV from top to bottom and the whole vocabulary of an IP fits on one screen. Its 1,500 headwords line up, showing their categories, worldview axes, and aliases. This is the first moment at which the full vocabulary map of the IP can be viewed at once.
Chapter 6 takes up worldview axes in earnest. It examines how the worldview-axis concept introduced here is designed for each IP and how it stands orthogonally to the nine categories to draw the Vault’s terrain.
© 2026 MEJE WORKS Corp. & 김동은WhtDrgon. All rights reserved.