dai 1.0.0.6
dotnet add package dai --version 1.0.0.6
NuGet\Install-Package dai -Version 1.0.0.6
<PackageReference Include="dai" Version="1.0.0.6" />
<PackageVersion Include="dai" Version="1.0.0.6" />
<PackageReference Include="dai" />
paket add dai --version 1.0.0.6
#r "nuget: dai, 1.0.0.6"
#:package dai@1.0.0.6
#addin nuget:?package=dai&version=1.0.0.6
#tool nuget:?package=dai&version=1.0.0.6
Dan — Document to AI-Friendly Markdown
Part of the RAG Document Toolkit for .NET, Version 1.0 — Sub Systems, Inc.
Convert DOCX, RTF, and HTML into structure-rich, RAG-ready chunks — with vector enrichment built in.
Dan converts DOCX, RTF, HTML, Markdown, and plain-text documents into token-sized Markdown chunks, each paired with a metadata record. The chunks are ready to embed and store in any vector database; the metadata lets your application filter, cite sources by document / section / page, and later hand the chunks to Ven (the companion Vector Enrichment Library) for structure-aware expansion.
| Namespace | SubSystems.RagDocumentToolkit.Dan |
| Assembly | DAN.DLL |
| NuGet package | dai |
| Minimum framework | .NET 9.0 |
| Companion library | Ven (vei package) |
Why Dan
Generic converters flatten a document into text and cut it at arbitrary character counts. Dan is built on the Sub Systems TE document engine, so it sees the document the way a word processor does and keeps that structure in the output:
- Real pagination and sections. Every chunk knows its page number and section, so answers can cite "Document, Section, Page" instead of a chunk number.
- Tables that survive chunking. Tables are emitted as HTML with row and cell ids, so a table split across chunks can be reassembled — by the model, or by Ven's table completion.
- Track changes and comments. Revisions and reviewer comments are carried into the Markdown with author information, and deleted text is explicitly marked so the model can ignore it unless asked.
- Self-qualifying chunks. Each chunk begins with a breadcrumb (document title, page, section, and so on), so it remains meaningful in a pool that mixes chunks from many documents and many providers.
- Token-accurate sizing. Chunk size is specified in tokens, not characters. Dan measures with the o200k encoding as it writes and adapts to the text, which matters when Latin text runs 3–5 characters per token and CJK text runs close to 1.
- Queryable metadata. Each chunk has a
Dictionary<string, object>record (values are strings and numbers only) covering document id, title, sequence, page, section, table / revision / comment information, chunk token count, and any key/value pairs you supply.
Installation
dotnet add package dai
Ven is optional but recommended; install the vei package to add vector enrichment and document-index creation.
Quick start
using SubSystems.RagDocumentToolkit.Dan;
// Once at program start: set your license information.
// See DasSetLicenseInfo in the help file for the parameters.
// Dan.DasSetLicenseInfo(...);
var dan = new Dan();
if (!dan.DasImportFile(@"C:\docs\service_agreement.docx", Dan.DOC_DOCX))
{
Console.WriteLine(dan.DasGetLastMessage());
return;
}
// Page numbers are zero-based; -1, -1 converts the entire document.
Dan.DanResult result = dan.DasGetMarkdown(-1, -1);
Console.WriteLine($"{result.DocTitle} ({result.DocId}): {result.chunks.Length} chunks");
for (int i = 0; i < result.chunks.Length; i++)
{
string chunkText = result.chunks[i]; // Markdown, breadcrumb first
Dictionary<string, object> meta = result.MetaRecs[i]; // metadata for this chunk
// Embed chunkText and store it, together with meta, in your vector database.
// Keep every meta record of the document: Ven needs the full set later.
}
dan.DasDispose();
When you send retrieved chunks to the model, use Dan.DasGetSystemMessage() as the system message. It explains the chunk format to the model — breadcrumbs, table ids, revision and comment tags — and how to cite sources.
Input formats
| Constant | Format | Import methods |
|---|---|---|
DOC_DOCX |
Word DOCX | DasImportFile, DasImportDocFromBytes |
DOC_RTF |
Rich Text Format | DasImportFile, DasImportDocFromString |
DOC_HTML |
HTML | DasImportFile, DasImportDocFromString |
DOC_MD |
Markdown | DasImportFile, DasImportDocFromString |
DOC_TEXT |
Plain text | DasImportFile, DasImportDocFromString |
PDF and Excel input are planned for a future release.
API at a glance
Import — DasImportFile, DasImportDocFromString, DasImportDocFromBytes
Convert — DasGetMarkdown(FirstPage, LastPage) returns a DanResult holding DocId, DocTitle, chunks, and MetaRecs.
Identify — Set the DocId and DocTitle properties before conversion to supply your own values. Otherwise Dan generates a document id and takes the title from the DOCX/RTF document information or the HTML title, falling back to the file path. Both properties are cleared after each DasGetMarkdown call, so a reused Dan object never carries one file's identity into the next.
Customize — DasOverrideChunkFieldName renames metadata fields to match the naming used by your other chunk providers. (The ssDocId and ssSeq fields are fixed; Ven relies on them.) DasSetFlags sets operation flags.
Events — LogMsg reports progress and diagnostic messages. MdoNames is called with the plain text of each chunk as conversion proceeds; your handler may return person names, organization names, place names, and important domain terms (from your own NER or keyword logic), which Dan adds to the chunk's metadata for pre-filtering.
Diagnostics and cleanup — DasGetLastMessage, DasResetLastMessage, DasDispose
Static — DasSetLicenseInfo, DasGetSystemMessage
The complete reference, with a six-step code walkthrough, is in the toolkit help file, rag_document_toolkit_help.htm.
Demo program
The toolkit includes a multi-file document Q&A demo in C# showing a complete pipeline: conversion with Dan, embedding and storage, AI-assisted document pre-selection, retrieval, enrichment with Ven, and answers that cite document, section, and page. In addition to the toolkit packages, the demo uses the LangChain 0.17.0 packages (Core, DocumentLoaders.Abstractions, Providers.OpenAI, Databases.InMemory, Databases.Sqlite) and the OpenAI .NET SDK. None of these are required by Dan itself.
Licensing
Dan is a commercial product licensed per developer, in Desktop, Server, Unlimited Server, and Enterprise editions. One license covers both Dan and Ven, and a single Dan.DasSetLicenseInfo call activates both libraries. The full license agreement is in the toolkit help file.
Support
Sub Systems, Inc. 3200 Maysilee Street, Austin, TX 78728 512-733-2525 https://www.subsystems.com
Copyright © 2026 Sub Systems, Inc. All rights reserved.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net9.0 is compatible. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net9.0
- LangChain (>= 0.17.0)
- LangChain.Core (>= 0.17.0)
- LangChain.DocumentLoaders.Abstractions (>= 0.17.0)
- LangChain.Providers.OpenAI (>= 0.17.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.0.0.6 | 69 | 10/4/2026 |