Bộ công cụ nghiên cứu AI

Hướng trợ lý AI của riêng bạn vào kho văn bản này — mô hình của bạn, tài khoản của bạn, chúng tôi không cần khóa nào. Kết nối qua MCP, hoặc gọi thẳng API công khai, để trợ lý duyệt, đọc, tìm và lập bảng ngữ cảnh 96,093 trang thuộc 750 văn bản.

Bạn không cần biết lập trình, cũng không cần cài đặt gì. Bước 1 bên dưới dùng được với bản miễn phí của hầu hết các trợ lý trò chuyện.

1

Bắt đầu ở đây — không cần cài đặt gì

Nếu bạn đã dùng ChatGPT, Claude, Gemini hay Copilot trên trình duyệt, bạn có thể đưa thư viện này vào việc chỉ trong khoảng một phút.

  1. Sao chép đoạn văn bản bên dưới bằng nút Sao chép.
  2. Mở một cuộc trò chuyện mới với trợ lý và dán vào làm tin nhắn đầu tiên. Có vẻ như không có gì xảy ra — điều đó bình thường. Đó là hướng dẫn, không phải câu hỏi.
  3. Bây giờ hãy đặt câu hỏi nghiên cứu bằng ngôn ngữ thường ngày. Trợ lý sẽ tìm trong thư viện này và trả lời kèm liên kết dẫn về trang gốc.
Lời nhắc trợ lý nghiên cứu
You are a research assistant for classical Vietnamese Hán-Nôm texts. You have read-only access to the Việt Điển corpus — 750 digitized Hán-Nôm texts (96,093 pages of OCR) — through a public JSON API (HTTP GET, no auth) at https://www.viet-dien.com.

Endpoints — reach for distribution and concordance first; they answer most
questions in one call, where search makes you page through results:
- https://www.viet-dien.com/api/v1/distribution?q=<query>          Which texts carry <query>, and how heavily, heaviest first. Call this first on a broad question, to choose what is worth reading.
- https://www.viet-dien.com/api/v1/concordance?q=<query>&width=30  Every occurrence of <query> across the corpus, one line each with context either side. This is how you read usage: senses, collocations and set phrases show themselves when the occurrences are lined up together, and one-snippet-per-page search hides them.
- https://www.viet-dien.com/api/v1/search?q=<query>               Pages whose OCR contains <query>, one snippet each. Returns text_id, title, title_han, item_id, page_number, snippet, viewer_url.
- https://www.viet-dien.com/api/v1/texts?q=<title>                Find texts by romanized or Hán (Sinitic) title.
- https://www.viet-dien.com/api/v1/texts/<id>                     A text's metadata and list of page numbers.
- https://www.viet-dien.com/api/v1/texts/<id>/pages/<n>?context=1  The OCR of one page, plus the source page-image URL. A classical sentence runs straight across the page break, so pass context=1 when quoting near a page edge, or the clause ends mid-air.
- https://www.viet-dien.com/api/v1/texts/<id>/full.txt            A whole work as plain text, pages marked — use it instead of forty page fetches when you need to read a text through.
- https://www.viet-dien.com/api/v1/variants?q=<char>              Which other glyph forms search treats as this character, so you can explain a match rather than assert it.

Search syntax (all of it works in q=, and is worth using — the words of a
classical collocation are often separated, so a plain phrase search misses them):
  國家              the characters together, in order
  "見聞 小錄"        quoted phrase
  國 AND 家         both somewhere on the same page
  國家 OR 天下       either one
  國 NOT 家         the first, without the second
  國 NEAR/10 家     within 10 characters of each other
  國?家             ? is any single character
  見聞*錄            * is a short gap (up to 8 characters)
Malformed queries return HTTP 400 with a "detail" message explaining the problem.

Requests are rate limited to 60/minute and 1000/hour per client, so work through
results steadily rather than fanning out; HTTP 429 means slow down and retry.

How to help me:
1. Search first, then fetch the specific pages you cite.
2. The material is Classical Chinese / Chữ Nôm and the OCR is machine-generated, so it may contain errors — read critically and flag uncertainty.
3. Always include Việt Điển links in your answers: cite every source as a clickable link to its viewer_url (a page on https://www.viet-dien.com) so I can open the original page image and OCR. Never cite a page without linking it.
4. For a character or phrase, tell me which texts and pages it appears on, link each with its viewer_url, and quote the surrounding snippet.

Corpus source: Digitizing Vietnam (https://www.digitizingvietnam.com).

Cách này hiệu quả nhất với trợ lý có thể truy cập web. Nếu trợ lý trả lời rằng nó không mở được liên kết, tức là nó không tới được thư viện — hãy thử bước 2, hoặc một trợ lý khác.

2

Kết nối qua MCP

không bắt buộc

Một số trợ lý có thể được trang bị công cụ riêng. Nếu trợ lý của bạn có mục Connectors, Integrations hay MCP servers, thì thêm một địa chỉ duy nhất bên dưới sẽ tốt hơn bước 1: trợ lý có được công cụ tra cứu chuyên dụng, và bạn không bao giờ phải dán hướng dẫn nữa.

URL máy chủ
https://www.viet-dien.com/mcp

Thêm vào đâu

  • Claude (claude.ai hoặc ứng dụng máy tính): Cài đặt → Connectors → Add custom connector, rồi dán địa chỉ vào. Có ở các gói trả phí.
  • ChatGPT: Cài đặt → Connectors, rồi thêm như một MCP connector tùy chỉnh. Có thể phải bật chế độ nhà phát triển trước.
  • Cursor, Claude Code và hầu hết công cụ khác có nhắc tới “MCP”: tìm mục Add server và chọn kiểu HTTP hoặc URL từ xa.

Chữ trong các menu này thỉnh thoảng thay đổi; điều bạn cần tìm là bất kỳ chỗ nào hỏi địa chỉ của một máy chủ MCP từ xa. Sau khi thêm xong, trợ lý đã biết cách dùng thư viện, nên bạn hỏi thẳng được — không phải dán gì cả.

Nâng cao: dán vào tệp cấu hình thay vì làm như trên
{
  "mcpServers": {
    "viet-dien": {
      "type": "http",
      "url": "https://www.viet-dien.com/mcp"
    }
  }
}
3

Nên hỏi gì

Cứ hỏi bằng ngôn ngữ thường ngày — tiếng Việt, tiếng Anh hay tiếng Trung — và đưa kèm chữ Hán nếu bạn có. Một vài câu hỏi hợp với kho tư liệu như thế này:

“Những văn bản nào ở đây bàn về 科舉, tức khoa cử, và văn bản nào nói nhiều nhất?”

Một nước mở đầu tốt cho bất kỳ chủ đề nào: nó cho biết tác phẩm nào đáng đọc trước khi bạn bỏ công đọc.

“Tìm mọi chỗ xuất hiện 皇越, và cho tôi xem dòng chữ quanh mỗi chỗ.”

Xếp các lần xuất hiện cạnh nhau là cách thấy được một cụm từ có nghĩa gì khi dùng thật, thay vì nghĩa từ điển.

“見聞 có xuất hiện gần 小錄 ở đâu không? Không nhất thiết liền nhau — trong khoảng mười chữ là được.”

Các chữ của một cụm từ cổ điển thường bị tách rời trên trang. Hỏi “gần” thay vì “đúng y” sẽ tìm ra điều mà tìm kiếm thường bỏ sót.

“Đọc mười trang đầu của Kiến văn tiểu lục quyển 1 và tóm tắt nội dung từng trang.”

Để định hướng trong một tác phẩm chưa quen, trước khi quyết định đọc kỹ chỗ nào.

Đọc câu trả lời thế nào

  • Văn bản mà trợ lý đọc là bản nhận dạng tự động từ ván khắc, chưa được hiệu đính. Có những chữ đơn giản là sai.
  • Hãy bấm vào liên kết. Mọi câu trả lời đều nên kèm liên kết tới các trang trên site này, nơi ảnh gốc nằm cạnh bản nhận dạng. Khẳng định nào không có liên kết thì coi như chưa được kiểm chứng.
  • “Không tìm thấy” là bằng chứng yếu. Một cụm từ có thể vẫn có mặt nhưng bị nhận dạng sai, hoặc nằm trong tác phẩm chưa được số hóa. Vắng mặt ở đây không có nghĩa là vắng mặt trong thư tịch.
  • Hãy trích dẫn bản gốc, không phải trợ lý. Theo liên kết, tự đọc trang đó, rồi trích dẫn văn bản và số trang như khi bạn làm việc với sách thư viện.
Dành cho lập trình viên Tài liệu API, danh sách công cụ, lược đồ

API công khai

Cùng kho văn bản đó qua HTTP GET thuần, dành cho ứng dụng không hỗ trợ MCP. JSON chỉ đọc, hỗ trợ CORS. Không cần khóa. URL gốc: https://www.viet-dien.com

GET /api/v1/stats Corpus totals Thử ↗
GET /api/v1/texts?q=易&limit=20 List / filter texts by title Thử ↗
GET /api/v1/texts/11 A text's metadata + page numbers Thử ↗
GET /api/v1/texts/11/pages/1?context=1 A page's OCR + image URL, with the pages either side Thử ↗
GET /api/v1/texts/11/full.txt A whole work as plain text Thử ↗
GET /api/v1/search?q=國&limit=20 Full-text OCR search Thử ↗
GET /api/v1/concordance?q=國家&width=30 Every occurrence of a term, in context Thử ↗
GET /api/v1/distribution?q=國家 Which texts carry a term, heaviest first Thử ↗
GET /api/v1/variants?q=為 A character's interchangeable forms Thử ↗
GET /api/openapi.json OpenAPI 3 schema (for tool import) Thử ↗

Những công cụ trợ lý của bạn sẽ có

search_corpus Find pages whose text contains a query, with a snippet of each. Good for "where does this appear"; for "how is it used" prefer concordance, and for "which works" prefer distribution.
concordance Every occurrence of a term across the corpus, one line each with context either side (a KWIC concordance). This is the tool for reading usage: senses, collocations and formulae show up when the occurrences are lined up together, which one-snippet-per-page search hides.
distribution Matching pages grouped by work, heaviest first. Call this first on any broad question: it tells you where a term lives before you spend calls reading pages.
get_page One page's OCR text, with the URL of the source page image and a viewer link to cite. Pass context to pull in the pages either side — a classical sentence runs across the page break, so a quotation from one page alone can end mid-clause.
get_full_text A whole work (or a run of pages) in one call, pages marked. Use it instead of forty get_page calls when you need to read a text through rather than check a citation. Ask for a page range on a long work.
list_texts The works in the corpus, filtered by title — romanized (Đại Việt sử ký) or Sinitic (大越史記) — or by catalogue field: author, year of composition, subject, language. Use it to find a work by who wrote it, when, or what it is about; use distribution to find one by content. Diacritics are ignored throughout, so "nguyen" finds "Nguyễn". Cataloguing is uneven: an author is recorded for about 39% of the corpus and a year for about 49%, so a year filter cannot see the undated remainder — do not report an empty or short result as evidence that the corpus holds nothing of that kind.
get_text_info One work's metadata and the page numbers actually digitized (they need not be contiguous), plus links to the viewer and to the item on Digitizing Vietnam.
char_variants Which other forms of a character search treats as the same. Woodblock carvers used whichever glyph they liked, so 為 and 爲 are one word; this reports the grouping, so you can explain a match rather than assert it.
corpus_stats How many works and pages are currently digitized. Worth checking before making a claim about coverage.

Các cách kết nối khác