Case Study: National Museum of Ethnology Reduces Workloads with Ondoku
Aug. 30, 2026
- National Museum of Ethnology
- Industry: Museum
- Location: Suita City, Osaka
- Interview: Mr. Kobayashi, Researcher, X-DiPLAS Project
Objectives / Challenges
Used for narration in video works related to cultural anthropology and ethnology. We used to have researchers active around the world record their voices, but recording could not be done well online, causing significant time and effort.
Solution
Abolish manual voice recording by researchers and create narration with Ondoku.
Effects
The need to tie up researchers' time for multiple recording sessions was eliminated, significantly reducing the workload. The speed of video production also improved, making it possible to handle tight schedules.
In this article, we introduce how Ondoku is being used at the National Museum of Ethnology (Minpaku) in Suita City, Osaka, as a case study for its implementation.
Introduction of the Organization and Department
Mr. Kobayashi (hereinafter, Kobayashi): The National Museum of Ethnology is a research institute for cultural anthropology and ethnology with museum functions. It is home to researchers who conduct fieldwork all over the world and provide their survey results widely to the public. It also has an affiliated "Graduate University for Advanced Studies" for cultural research, where students aiming to complete their doctoral dissertations deepen their learning every day. Our museum, affectionately known as "Minpaku," celebrated its 50th anniversary in 2024.
I have been working as a researcher for X-DiPLAS since the 2022 fiscal year. X-DiPLAS is a project that aims to create a database of photographs taken by cultural anthropologists and archaeologists around the world and build an environment where the public can freely view them. Currently, we are focusing on activities that reflect "Digital Stories"—showing what kind of story is behind each individual photo—rather than just storing the photos.
HP: National Museum of Ethnology
Please tell us about the background of introducing the text-to-speech tool.
Kobayashi: We introduced the text-to-speech tool primarily to reduce the burden on our researchers.
I continue to work on databasing photographs taken by cultural anthropologists and archaeologists from all over the world and making them available to the public. However, simply collecting photos is not enough to correctly inherit history. Every photo has a story, so it is necessary to produce video works that include the "voice" of the researcher explaining the situation at the time the photo was taken.
Initially, we contacted researchers around the world and recorded their voices for video works online. However, issues such as voice lag and external noise meant that we often had to re-record multiple times, which was a problem. While looking for something that could reduce the burden on researchers, I came across text-to-speech tools that can create a "voice" simply by entering text.
May we ask how you came to introduce Ondoku?
Kobayashi: I used to have a bit of a resistance to systems that use AI, such as text-to-speech tools. I felt a sense of discomfort with voices that seemed somewhat detached from reality. However, Ondoku’s voices are natural, and I felt they were perfect for replacing the researchers' narration.
Additionally, the excellent cost-performance was a factor in our evaluation. The balance between ease of use and pricing is good, and it performs exceptionally well even within our limited budget.
Have the challenges been resolved (improved) compared to before introducing Ondoku?
Kobayashi: Not only has it reduced the burden on researchers, but we are also now able to produce video works more smoothly.
The recording work, which tied up the researchers' time, was one of the negative aspects of our work. Thanks to Ondoku, communication with researchers has improved, which is a major step forward. Furthermore, high-quality audio is completed just by typing text, which contributes to increasing the speed of video production.
We regularly hold symposia to promote our projects. Using Ondoku, we were able to smoothly start creating the necessary materials for presentations. I feel the wide range of Ondoku’s effectiveness particularly when preparation time for a symposium is tight. In fact, during periods when production is busy, I use Ondoku almost every day.
In addition to its use in research, please tell us about any other cases where Ondoku was helpful.
Kobayashi: The introduction of Ondoku has improved the quality of the scripts written for narration.
When writing, there are times when you accidentally add unnecessary expressions and sentences become too long. Actually, if you input an unnecessarily long sentence into Ondoku, it doesn't read it back very well. Because it speaks like a human, there are times when you notice a sense of discomfort in the timing of breaths.
When I rely on the sound, review the expressions, and try again, I realize that I have created a higher-quality script. The ability to create concise and clear writing, rather than just reading out audio, is one of the charms of Ondoku.
If you have any further requests for improvements to Ondoku, please let us know.
Kobayashi: From the perspective of a researcher working in the field overseas, I would be happy to see the addition of Swahili. Swahili is a language with a wide range of use within Africa, so having it in the lineup would broaden the scope of our work.
Of course, for even smoother utilization, it might be good to have system improvements like intonation control. However, if it becomes too functional, the simple operability might be lost. Considering the balance including the price, I feel that the current interface is just the right level.
How do you plan to use Ondoku in the future?
Kobayashi: I would like to increase the opportunities to use Ondoku in educational activities as well.
I sometimes give assignments to students at the university affiliated with the museum to create short video works. During assignment presentations, I once recommended the introduction of narration by Ondoku as a means to improve the quality of the work. I feel that since Ondoku is easy for even students to use, there are many situations where it could be useful in class.
Also, I would like to leverage Ondoku's strength in foreign language support for future activities. We currently have plans to create new video works using English narrations based on English translations created by overseas researchers.
If we were to outsource English voice recording, no amount of budget would be enough. With Ondoku, it should contribute to overwhelming cost savings. I want to continue the activity of capturing the voices of researchers by utilizing Ondoku.
By utilizing Ondoku, you were able to not only reduce the burden on researchers for voice recording but also reduce the time-related costs of producing video works! Thank you for sharing such a wonderful case study.
■ AI voice synthesis software "Ondoku"
"Ondoku" is an online text-to-speech tool that can be used with no initial costs.
- Supports 81 languages and regions, including Japanese, English, Chinese, Korean, Spanish, French, and German
- Available from both PC and smartphone
- Suitable for business, education, entertainment, etc.
- No installation required, can be used immediately from your browser
- Supports reading from images
To use it, simply enter text or upload a file on the site. A natural-sounding audio file will be generated within seconds. You can use voice synthesis up to 5,000 characters for free, so please give it a try.
Email: ondoku3.com@gmail.com
"Ondoku" is a Text-to-Speech service that anyone can use for free without installation. If you register for free, you can get up to 5000 characters for free each month. Register now for free
- What is Ondoku
- Start text-to-speech conversion
- Free registration
- Pricing
- Posts
- Try other free services
