Online Kanbun tool

331 views
Skip to first unread message

Siva Kalyan

unread,
Sep 13, 2026, 10:16:14 AMSep 13
to pm...@googlegroups.com
Hi all,

I thought the members of this list might be interested in an online tool I created, that lets you enter a text in Literary Chinese, and automatically generates both the kundokubun and the kakikudashibun: https://jidou-kundoku.vercel.app/

The underlying syntactic analysis is directly editable, so you can correct any mistakes in the generated kundoku; corrections are immediately reflected in the kaeriten as well as the kakikudashibun.

You can individually toggle the display of furigana, okurigana, and kaeriten (or turn them all off if you just want the hakubun), and you can print the text with whichever annotations you choose.

Finally, poetry is automatically annotated with the Qièyùn rime category of the final character of each line.

Please feel free to inform me of any errors (of which I’m sure there are many), and to let me know if there are any other features you would like to see.

I hope this tool is of some use to scholars of pre-modern Japan.

Regards,
Siva

Laffin, Christina

unread,
Sep 13, 2026, 3:50:25 PMSep 13
to pm...@googlegroups.com

Dear colleagues,

While I welcome greater interest in and resources for the study of kanbun in its various forms, I wonder if this might not be a good time to remind ourselves about the wealth of variation across history and region in terms of practices of kanbun kundoku and the excellent scholarship by so many on this.

In a moment when the value of our humanistic knowledge is being challenged and the need for critical expertise and transparency is needed more than ever, I worry about how this approach can flatten our understanding and engagement with kanbun. I wish to know more about the underlying expertise and process in developing this tool, the functions for which it is suitable, what work and scholarship it is built on, etc. In particular, because the framework appears to be built on Claude or a similar system, this also brings to mind ethical questions about not only scholarship, but land, labour, and lack of transparency around impact on our planet. Many of us are currently awaiting our claims from the Anthropic Copyright Settlement after our scholarly works, the culmination of many years of research, were stolen.

Undoubtedly, we are in a moment when it is much easier for these tools to be created, but I wish for and encourage more openness about the how, why, and for whom these tools are being made.

Christina

Christina Laffin クリスティーナ・ラフィン (she, her, hers)
Associate Professor, Department of Asian Studies
The University of British Columbia | x
ʷməθkʷəy̓əm (Musqueam) Traditional, Ancestral, Unceded Territory

 

From: pm...@googlegroups.com <pm...@googlegroups.com> On Behalf Of Siva Kalyan
Sent: Sunday, September 13, 2026 1:31 AM
To: pm...@googlegroups.com
Subject: [PMJS] Online Kanbun tool

 

[CAUTION: Non-UBC Email]

--
PMJS is a forum dedicated to the study of premodern Japan.
To post to the list, email pm...@googlegroups.com
For the PMJS Terms of Use and more resources, please visit www.pmjs.org.
Contact the moderation team at mod...@pmjs.org
---
You received this message because you are subscribed to the Google Groups "PMJS: Listserv" group.
To unsubscribe from this group and stop receiving emails from it, send an email to pmjs+uns...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/pmjs/D7C86B4A-CCC9-4000-A009-85C087A7CEF4%40gmail.com.

Siva Kalyan

unread,
Sep 15, 2026, 7:16:06 PMSep 15
to pm...@googlegroups.com
Dear Prof. Laffin (and all),

Thank you for raising these important issues. Let me use the opportunity to provide some clarification regarding the background and aims behind the tool.

The tool's architecture and interface were indeed designed with the assistance of Claude (with many rounds of manual verification against printed editions of kanbun texts to ensure that the typography is accurate). However, the functioning of the tool itself does not use LLMs in any way. It is performed using a dependency parser that I trained from scratch on a version of the open-source Classical Chinese Universal Dependencies Treebank, which was manually annotated at the Instiitute for Research in Humanities at Kyoto University, and contains the full text of works such as 論語, 孟子, and 礼記, taken from the Kanseki repository.

The output of the dependency parser is converted into kundoku using a system of rules that I initially curated by hand (from my limited knowledge of kanbun), but later developed through comparison with the readings at https://kanbun.info, which I developed into an automated test suite. The character readings were taken from a combination of Wiktionary (via the dumps at https://kaikki.org) and NINJAL’s UniDic. Rime categories for poetry were taken from the qieyun Python package. Kakikudashibun is derived from the kundoku in the purely mechanical way that we are all familiar with.

Finally, a word about the tool’s aims. This is a fully open-source project (source code available here), intended as a convenience for researchers accustomed to accessing Literary Chinese texts via kanbun. It is not intended as a replacement for consulting the many excellent printed kanbun editions of Chinese texts. The application does not (and cannot) collect user data; once the page has loaded, it will work even without an Internet connection (all of the syntactic analysis and transformation is done on the user’s device).

One of the things I came to learn through the process of building this tool is precisely the many subtle layers of variation in the choice of readings (including whether to end a 連用形 form with て, where and when to use は, etc.), and I am acutely conscious of the fact that over-reliance on such a tool by researchers or (particularly) students could lead to the perpetuation of a flattened and ahistorical set of conventions. However, I think (and hope) that such a tool will have enough practical use to at least partly outweigh these concerns 

Regards,
Siva

Mikhail Skovoronskikh

unread,
Sep 15, 2026, 10:35:52 PMSep 15
to pm...@googlegroups.com
Dear Siva,

I am sorry if this sounds like a silly question, but is the main purpose of the tool to quickly generate a punctuated kunten version of a text that can be easily edited and exported into other applications and digital environments? As you know, kanbun editing in Word is a nightmare, and there are Japanese software products that claim to make the process easier. Or is the tool's primary purpose to produce reliable kakikudashi renderings after all?

Best regards,
Mikhail Skovoronskikh




Ross Bender

unread,
Sep 15, 2026, 10:36:33 PMSep 15
to pm...@googlegroups.com
Dear all,

Thanks to Siva Kalyan for his detailed explanation of how he developed this tool. I tried it out, and it was rather ingenious, but it had some limitations. For example, here it rendered Daigokuden (the Imperial Council Hall) as "Oi nari Gokuden," which is misleading.

image.png

In the following example, the tool completely missed the name of the pavilion --   "Inokuma -- 猪熊院.   , My translation of the passage: "The Emperor went to view several pavilions and arrived at the Inokuma Pavilion where officials of fifth rank and up held archery contests. He gave cash gifts to the archers. After the contests he gave silk to officials of fifth rank and up and noblewomen according to status."

image.png

However, the tool is rather ingenious and with time and corrections could prove to be of some value. I would applaud such efforts, rather than discourage them.

Respectfully, 
Ross Bender


On Tue, Sep 15, 2026 at 7:16 PM Siva Kalyan <sivakalyan...@gmail.com> wrote:

GUELBERG Niels

unread,
Sep 16, 2026, 10:17:51 AMSep 16
to pm...@googlegroups.com
Why are you getting so upset, Christina?

I see it as a chance to get young people interested in kanbun. The common approach (to let the students learn the ahistorical yomikudashi-system established by the Meiji government) doesn't make sense except for the few one who want to pass an admission exam for kokugo at a Japanese university

I tested the app with the preface of the Zhoubi suanjing (周髀算經) because it is one of the texts which were definitely used in Japan in the midth of the 13th century.

Zhoubi suanjing:
夫高而大者,莫大於天;
厚而廣者,莫廣於地。
體恢洪而廓落,形脩廣而幽清,
可以玄象課其進退,然而宏遠不可指掌也。
可以晷儀驗其長短,然其巨闊不可度量也。
雖窮神知化不能極其妙,探𦣱索隱不能盡其微,
是以詭異之說出,則兩端之理生,遂有渾天、蓋天,兼而並之。

故能彌綸天地之道,
有以見天地之𦣱,
則渾天有靈憲之文,蓋天有周髀之法,
累代存之,官司是掌,
所以欽若昊天,恭授民時。

爽以暗蔽,才學淺昧,
隣高山之仰止,慕景行之軌轍,負薪餘日,聊觀《周髀》。
其旨約而遠,其言曲而中,將恐廢替,濡滯不通,使談天者無所取則,
輒依經為圖,誠冀頹毀重仞之墻,披露堂室之奧,庶博物君子,時逈思焉。

First I did what most of my students would do: let the work do by ChatGPT.

chatgpt:
夫れ高くして大なる者は、天より大なるは莫く、厚くして廣き者は、地より廣きは莫し。
體は恢洪にして廓落、形は脩廣にして幽清なり。
以て玄象を課して其の進退を知るべし、然れども其の宏遠なること、指掌すべからざるなり。
以て晷儀をもって其の長短を驗すべし、然れども其の巨闊なること、度量すべからざるなり。

神を窮め化を知ると雖も、其の妙を極むること能はず、
𦣱を探り隱を索むると雖も、其の微を盡くすこと能はず。
是を以て詭異の說出づれば、則ち兩端の理生じ、遂に渾天・蓋天有り、兼ねて之を並ぶ。

故に能く天地の道を彌綸し、以て天地の𦣱を見ること有れば、
則ち渾天には『靈憲』の文有り、蓋天には『周髀』の法有り。
累代之を存して、官司是れを掌る。
所以に昊天を欽若し、恭しく民の時を授く。

爽、暗蔽にして、才學淺昧なり。
高山の仰止を隣み、景行の軌轍を慕ひ、薪を負ふ餘日、聊か『周髀』を觀る。
其の旨は約にして遠く、其の言は曲にして中たる。
將に廢替し、濡滯して通ぜざらんことを恐れ、天を談ずる者をして取則する所無からしめんとす。
輒ち經に依りて圖を為り、誠に頹毀せる重仞の墻を重ねて、堂室の奧を披露せんことを冀ふ。
庶はくは博物の君子、時に逈かに之を思はんことを。

And this is the result of the app:

jidou-kundoku:
夫れ高くて大いなる者は、天より大いなる莫(!)、厚しくて廣る者は、地より廣る莫(!)。
體洪を恢して廓落つること、形廣りて幽しこと清きを脩め、
玄象を以て其の進退するを課す可し、然りて宏くわう遠し掌を指す可からざるなり。
以て儀ぎを晷き其の短きを長きを驗ぐ可し、其の巨こと闊し然りて量を度い可からざるなり。
神を窮むと雖も化くるを知ること其の妙を極むること能はず、
𦣱探して索隱すること其の微かをず盡くすこと能はざるに、
是を詭き異いての說を以て出づ、則ち兩端の理生にして、渾こん天てんを遂に有り、天を蓋ひ、兼ねて之を並ぶ。

故に能く天地の道を彌み綸りん、以て天地の𦣱を見る有れば、
則ち渾こんの天に靈れい憲けんするの文有り、
蓋し天に周しうの髀の法有り、
累るゐ代だいに之を存り、官くわん司し是の掌なり、
以て所昊の天を欽きん若じやくにして、恭しく民に時を授かる。

爽らかは暗あん蔽へいて以て、才は學は淺せん昧まいし、
高山の仰ぎやう止しするを鄰り、景けいを慕ぼにして之に軌の轍を行ふ、薪を負かすこと餘よ日にちに、聊《周しう髀ひ》を觀る。
其の旨きこと、其の言ふは曲がりて中り、
將に替わる廢するを恐れんとす、濡じゆ滯たいすること通らず、天を談する者取る所則ち無からしめ、
すなはち經けいを依り圖を為し、冀頹を誠重じゆう仞じんの墻を毀ち、露の堂の室の奧を披し、
庶しよ博はくて君子を物、時に逈て思ふ約めて遠し。

It is obviously not an LLM but a little program which implemented some of the rules of schoolbook kanbun. But it may be useful to print kanbun sheets (I didn't checked it).

The app needs some improvements and by improving it young people can learn more about some issues in kanbun grammar. And at some point they may reach the conclusion that you have to understand and interpret the text first before you start to translate it.

A more straight way to get to that point is (for non-Japanese students) to do your yomikudashi in English (or in German, as in my case). I was often asked by my Japanese colleagues why I am so fast in reading kanbun texts. The answer is: it is far more easy to transform a chinese syntax sentence in an European language than in Japanese.

For 35 years I spent most of my time in deciphering and editing kanbun texts. For a few of them I even provided a yomikudashi-version for my Japanese readers. But it is very difficult to arrive at a historically grounded reading. Indeed there exist excellent scholarship and I had the pleasure to get advices from the late Tsukishima Hiroshi and from Kobayashi Yoshinori (they invited me to join their research at Kozanji in the early 90th), but for me this work is the top rung of the ladder. Why should young people start from the top?

Niels



差出人: pm...@googlegroups.com <pm...@googlegroups.com> が Laffin, Christina <Christin...@ubc.ca> の代理で送信
送信日時: 2026年9月14日 4:34
宛先: pm...@googlegroups.com <pm...@googlegroups.com>
件名: RE: [PMJS] Online Kanbun tool
 

Paula R. Curtis

unread,
Sep 16, 2026, 10:35:25 AMSep 16
to pm...@googlegroups.com
Dear all,

We are all well aware that AI, LLMs, and the current state of publishing is a sensitive topic, one which elicits a wide variety of opinions across academic circles. This is a reminder that PMJS is a mailing list for scholarly discussion, and per our Terms of Use, responses should be first and foremost collegial, as well as grounded in scholarly evaluation. The Moderation Team does their best to catch any disrespectful or irrelevant commentary, though as individuals with their own busy inboxes, can miss some posts. Please be mindful to focus on constructive comments and treat colleagues appropriately.

Best,

Paula
PMJS Logistics Manager



--
Dr. Paula R. Curtis
Operations Leader, Japan Past & Present
Yanai Initiative for Globalizing Japanese Humanities
Academic Administrator
Department of Asian Languages & Cultures, UCLA

Portfolio & Projects

Michael McCarty

unread,
Sep 16, 2026, 11:24:11 AMSep 16
to pm...@googlegroups.com
Dear all, 

I do not know much about the technical issues of how these tools are made (so I definitely appreciate Siva's detailed explanation), nor do I teach at a university where I have the ability to instruct graduate students in kanbun kundoku. So in some ways I do not have much of a horse in this race. But knowing how often my own undergraduate students use various LLMs and tools to attempt to cheat, I do think we as educators need to be aware of the pedagogical downsides to making such processes "easier" for students, when the training of the brain is what is important. I'm very interested to hear what others who teach kanbun think.

And, respectfully, I'd like to add to Prof. Guelberg that from my perspective, Christina already explained her concerns fully and cogently, and nowhere did she sound "upset." 

Best,

Michael McCarty

Siva Kalyan

unread,
Sep 22, 2026, 9:28:15 PM (10 days ago) Sep 22
to pm...@googlegroups.com
Dear all,

Thank you all for trying out my tool (link here, for convenience). I have incorporated your feedback as best I can, and intend to continue making improvements to it—so please continue to try it out and send feedback!

One thing I may not have emphasised sufficiently in my original message is that the quality of the kundoku is limited by the accuracy of the parsing (〜品詞分解). Currently, the automated parsing is about 75% accurate, so most texts will require some manual adjustment. You can do this in the following ways:
  1. If you double-click or right-click on a character, you will see
    1. a label showing the character’s part of speech (e.g. 名詞). Hovering over the label pops up two further labels with semantic categories (e.g. 人). Clicking on any of these labels brings up a menu that lets you change the label.
    2. a dependency arrow leading to the character from its head or governor (e.g. from a verb to its object). You can reassign the head by dragging the character onto a different head.
    3. a label next to the dependency arrow showing the type of dependency (e.g. 主語, 目的語). Clicking on this label brings up a menu that lets you change the type of dependency. (Unlikely choices are dimmed, to make the choice less overwhelming.)
  2. If you right-click or double-click on furigana, a menu will pop up that lets you choose a different reading (consistent with the chosen part of speech).
Changes to the part of speech or semantic category will be immediately reflected in the furigana, and changes in the dependencies will be immediately reflected in the kaeriten and kakikudashibun. (All of this information can be found by clicking the "How to edit" button in the app.)

Addressing individual messages:

In response to Mikhail: the purpose of the tool is to produce a reasonable first draft of a kunten version, which can then be edited in place (as described above), and exported as either PDF, LaTeX, or CoNLL-U (a text-based tabular format common in NLP research).

Thanks to Ross for trying a passage from the 日本後紀. This has made me realise that it's necessary to point out that the app is not really designed for 和化漢文, as the parser was trained exclusively on texts from China. I have made a few improvements, though, so that at least a hand-edited parse will read OK.

Thanks very much to Niels for trying out a passage from 周髀算經, and for providing the (presumably satisfactory) ouput from ChatGPT. This has led to massive, noticeable improvements in the correctness of readings (including the instances of 莫 that he flagged), and the passage should read much better now (especially with a hand-corrected parse).

In response to Michael: I completely agree that making things easier for students is a double-edged sword. The best response I can give is that the tool’s accuracy is still sufficiently low that correcting its output is itself a good exercise for students (at least, I have found it to be so for myself).

Looking forward to more feedback; I hope that, as Ross said, with time and corrections it can prove to be of some value.

Regards,
Siva

Cynthea Bogel

unread,
Sep 23, 2026, 7:31:22 AM (9 days ago) Sep 23
to pmjs digesti subscribers

Dear colleagues,

Some of us began our academic careers with the remarkable "Sweet JAM," a program developed at MIT that let us type Japanese characters on our department's computers (unless we were lucky enough to have one of our own). We learned what one-byte and two-byte meant. We typed, printed our texts, peeled the perforated edges from the paper, and delighted in it all. Some lamented what we might lose by no longer writing characters by hand. That has come to pass, but not for everyone. One could argue that interest in calligraphy has grown precisely because we no longer write characters every day.

Cynthea

Cynthea J. Bogel



On Sep 22, 2026, at 9:28 PM, Siva Kalyan <sivakalyan...@gmail.com> wrote:

Dear all,

Gian Piero Persiani

unread,
Sep 23, 2026, 11:23:19 AM (9 days ago) Sep 23
to pm...@googlegroups.com

Dear All,

Cynthea makes a valid point. On the other hand, when Sweet JAM was released the value of academia was not being so bitterly disputed and the very survival of academia—and of the humanities in particular—was not in doubt. I think we need to exercise a great more deal of caution with technology now that the tech industry has the political and cultural sectors of society in its grip. In that regard at least, technology today is an entirely different beast.

Personally, I am not optimistic about the survival of philology and other disciplines that require years of training if too many buy into the myth of fast “knowledge” for everyone. Traditional time-honed scholarly expertise (i.e., academia) and “speed-of-output-first, worry-about-quality-later” tools can hardly be said to be compatible if the unstated goal of the latter is to make the former obsolete and unnecessary. So I for one completely understand the hesitation some feel in endorsing these tools and the world they may be ushering in.

Kind regards,

Gian Piero Persiani


Ross Bender

unread,
Sep 23, 2026, 5:53:31 PM (9 days ago) Sep 23
to pm...@googlegroups.com
Dear all,

Obviously this discussion is very timely, as AI tools have developed so rapidly. I, for one, worry that AI would deprive me of a job, if I had a job.

However, it seems to me that the potential doom of the humanities or philology in particular has less to do immediately with the growth of technology than that kids these days are more interested in careers in law, business, or consultancy, than in the good old disciplines of history or literature. (I speak from my observations of undergrads at University of Pennsylvania.)

As Cynthea Bogel observed, " Some lamented what we might lose by no longer writing characters by hand." This takes me back to my grad student days in the 70s at Columbia when we were required to hand write the kanji to all the Japanese and Chinese terms used in our dissertations. It also takes me back to the process of typing that thesis by hand on a Royal portable, using gallons of white out, when every slip in footnoting resulted in multiple and frustrating retypings. When that draft was finished, a very helpful secretary in the department produced a final typescript for a fee.

At that time we used the paperback copy of Tsuchihashi to transpose Japanese dates into those of the Julian calendar. It was very laborious. Now I use Nengo-Calc from University of Tuebigen. For many years now I have used Muller's excellent Digital Dictionary of Buddhism and his CKJV-E dictionary. Some years ago I acquired a Japanese word processor and dictionary called NJ Star. 

For many years I tracked the progress of Google Translate, which at first could render passable English translations of French and German, but not modern Japanese. Now it can do excellent modern Japanese, but not classical Chinese. However just in the past year I have discovered that ChatGPT can do a creditable rendition of literary Sinitic, although It obviously takes the skill of a long-suffering and trained philologist (me) to correct its errors and resist hallucinations.

Now I wonder whether going back to the error of laboriously carving wedges in clay tablets by the light of the moon would somehow be more virtuous. Also, it's far more likely apparently that AI will kill off humanity before it even finishes killing off the humanities. 

However, I am not in the position of having to teach university students to read Old Japanese or any other type of Japanese, and I sympathize with those in that position. 
On Wed, Sep 23, 2026 at 11:23 AM Gian Piero Persiani <gper...@gmail.com> wrote:

Dear All,

Cynthea makes a valid point. On the other hand, when Sweet JAM was released the value of academia was not being so bitterly disputed and the very survival of academia—and of the humanities in particular—was not in doubt. I think we need to exercise a great more deal of caution with technology now that the tech industry has the political and cultural sectors of society in its grip. In that regard at least, technology today is an entirely different beast.

Personally, I am not optimistic about the survival of philology and other disciplines that require years of training if too many buy into the myth of fast “knowledge” for everyone. Traditional time-honed scholarly expertise (i.e., academia) and “speed-of-output-first, worry-about-quality-later” tools can hardly be said to be compatible if the unstated goal of the latter is to make the former obsolete and unnecessary. So I for one completely understand the hesitation some feel in endorsing these tools and the world they may be ushering in.

Kind regards,

Gian Piero Persiani


On Wed, Sep 23, 2026 at 6:31 AM Cynthea Bogel <cjb...@gmail.com> wrote:

Dear colleagues,

Some of us began our academic careers with the remarkable "Sweet JAM," a program developed at MIT that let us type Japanese characters on our department's computers (unless we were lucky enough to have one of our own). We learned what one-byte and two-byte meant. We typed, printed our texts, peeled the perforated edges from the paper, and delighted in it all. Some lamented what we might lose by no longer writing characters by hand. That has come to pass, but not for everyone. One could argue that interest in calligraphy has grown precisely because we no longer write characters every day.

Cynthea

Cynthea J. Bogel


.

Iyanaga Nobumi

unread,
Sep 23, 2026, 10:42:08 PM (9 days ago) Sep 23
to pm...@googlegroups.com
Dear All,

Somehow related to this topic, recent news can draw out attention: that huge numbers of second-hand books are bought by some unknown buyers, and are disappearing from the old book stores' market. See for example <https://news.yahoo.co.jp/articles/b5bf35aac135601dde27ef7d43cbf5d7fe95e671>.

News sites doubt that the buyers are American AI companies, and these books are bought to be "eaten" by AI engines. Most books are from specialized fields of research.

We know that there are Web sites such as Internet Archive, NLD's Digital Library, and others, where it is possible to find and often download copyright-free old books, and these are of tremendous value for our research activities.The full text search capabilities of NLD are incredible tools of research. However, what this recent news let us know is much more disquieting. It may be that books still under protection of copyrights are taken out, and the last copies of books are "eaten"; what they had inside are broken into tiny pieces of electronic information, while as wholes, they may be completely thrown away...

All this is simply guesses, and nobody seems to know what is really happening. But I think we should be at least aware of the facts and try to have more accurate information, as well as reflect on this issue.

Best regards,

Nobumi Iyanaga
> --
> PMJS is a forum dedicated to the study of premodern Japan.
> To post to the list, email pm...@googlegroups.com
> For the PMJS Terms of Use and more resources, please visit www.pmjs.org.
> Contact the moderation team at mod...@pmjs.org
> ---
> You received this message because you are subscribed to the Google Groups "PMJS: Listserv" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to pmjs+uns...@googlegroups.com.
> To view this discussion visit https://groups.google.com/d/msgid/pmjs/CAMEQgpEwj%2Bqf0ozwSfCrc%3D%2BXMFMDE-aKxGD0S4CpuuHPVWN2sw%40mail.gmail.com.

Reply all
Reply to author
Forward
0 new messages