A technical appendix to Finding Jikishinkage-ryū Swordsmen in the Bakufu Registers and The Zoku Tokugawa Jikki and the Kōbusho. It assumes a command line and some patience; neither article depends on it.
Searching the sources
Looking at a page
Every frame in the National Diet Library’s collection has a stable address and can be fetched as an image:
# frame 82 of pid 2547156, 2400 pixels wide curl -o k0082.jpg \ "https://dl.ndl.go.jp/api/iiif/2547156/R0000082/full/2400,/0/default.jpg"
If you fetch more than a page or two, wait several seconds between requests and run one job at a time. The pages used here were fetched five seconds apart.
Searching the library’s own text
For printed books in movable type, such as Rikugun rekishi, the library has machine-read text, and its site searches inside it. The text of a volume can be downloaded from the National Diet Library Digital Collections (opens in a new tab) as a zip archive holding one small text file per frame, named <pid>_<frame>.txt, alongside a .json file that gives the position of each line on the page.
# every line that contains the name, across the whole volume unzip -p 988906.zip '*.txt' | grep '榊原鍵吉' # which frames contain it for f in $(unzip -Z1 988906.zip '*.txt'); do unzip -p 988906.zip "$f" | tr -d '\n' | grep -q '榊原鍵吉' && echo "$f" done
Two habits matter more than the commands.
Join the lines first. The text follows the printed columns, and a name that runs over the end of a column is split across two lines. tr -d '\n' removes the breaks so that the name is whole again. In rosters such as these the characters of a name are also spaced out to fill the column, so allow for gaps:
unzip -p 988906.zip '*.txt' | tr -d '\n' \
| grep -oP '.{0,15}男\s*谷\s*精\s*一\s*郎.{0,15}'
Search for the variants. Old and new character forms, and the machine’s misreadings, all have to be tried: 劍, 劔 and 剣; 將 and 将; 鍵 and 健; 總 and 総. A character class does this in one pass, and searching on the surname alone, then reading every hit, is slower but misses less. In books printed before the script reforms the old forms are the ones in the text, and a search on the modern form simply fails: 奧詰 and not 奥詰, 遊擊隊 and not 遊撃隊, 郞 and not 郎.
unzip -p 988906.zip '*.txt' | tr -d '\n' \
| grep -oP '.{0,10}天\s*野\s*[將将]\s*[曹曺].{0,10}'
The text is a finding aid. Every hit used above was then read on the page image, and several readings changed when it was.
Fetching the text of one volume
For open-access books the library’s research service, NDL Lab (opens in a new tab), returns the machine text of a volume page by page, two hundred frames to a request. The three volumes of the Zoku Tokugawa jikki used here came this way:
# frames 1-200 of pid 773008; repeat with from=200, 400, ... until the list is empty curl -o p-0.json \ "https://lab.ndl.go.jp/dl/api/page/search?f-book=773008&size=200&from=0"
Each item in the reply has the frame number (page) and its text (contents), which can be joined and searched as above. The same service searches the full text of the whole open collection for a phrase:
curl "https://lab.ndl.go.jp/dl/api/book/search?keyword=奧詰銃隊&size=40"
It covers only works that are out of copyright, and it treats the keyword as one phrase. As with the page images, wait several seconds between requests.
Reading a page by machine
The woodblock registers have no text layer worth the name, and the pages have to be read by eye or by a model you run yourself. Two free ones are worth knowing.
The library’s own model for premodern books, NDL Koten OCR-Lite (opens in a new tab), runs on an ordinary computer:
git clone https://github.com/ndl-lab/ndlkotenocr-lite cd ndlkotenocr-lite && python3 -m venv venv && source venv/bin/activate pip install -r requirements.txt && cd src python3 ocr.py --sourceimg k0082.jpg --output out # or --sourcedir <folder>
It writes a .txt, a .json and an .xml for each image, which can be searched with the commands above. For printed Meiji and later books, Tesseract (opens in a new tab) with its vertical Japanese model does well:
mkdir tessdata && cd tessdata curl -L -O https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata curl -L -O https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn_vert.traineddata cd .. TESSDATA_PREFIX=./tessdata tesseract k0127.jpg stdout -l jpn_vert --psm 5
How far to trust them: on the movable type of Rikugun rekishi either one gives text good enough to search. On the cursive of a bukan the results are poor. I ran the library’s model over all fifty-two openings of the 1863 pocket edition; it read the headings of the Gakumonjo and the Kaiseijo and little else, and it did not find the Kōbusho. The models are best used to decide which pages to read, and an AI assistant that can look at images is useful in the same way, to propose a reading that you then check against the page. Where a reading here is uncertain I have said so, and the links let you look for yourself.
