Unhwp 0.18.0

dotnet add package Unhwp --version 0.18.0
                    
NuGet\Install-Package Unhwp -Version 0.18.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Unhwp" Version="0.18.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Unhwp" Version="0.18.0" />
                    
Directory.Packages.props
<PackageReference Include="Unhwp" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Unhwp --version 0.18.0
                    
#r "nuget: Unhwp, 0.18.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Unhwp@0.18.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Unhwp&version=0.18.0
                    
Install as a Cake Addin
#tool nuget:?package=Unhwp&version=0.18.0
                    
Install as a Cake Tool

Unhwp

High-performance .NET library for extracting HWP/HWPX Korean word processor documents to Markdown.

Installation

dotnet add package Unhwp

Or via NuGet Package Manager:

Install-Package Unhwp

Quick Start

using Unhwp;

// Parse a document
using var doc = UnhwpDocument.ParseFile("document.hwp");

// Convert to Markdown
string markdown = doc.ToMarkdown();
Console.WriteLine(markdown);

// Plain text and JSON
string text = doc.ToText();
string json = doc.ToJson();

Console.WriteLine($"Sections: {doc.SectionCount}");

With Markdown Options

using var doc = UnhwpDocument.ParseFile("document.hwpx");

var markdown = doc.ToMarkdown(new MarkdownOptions
{
    IncludeFrontmatter = true,
    Refine = true
});

Parse from Bytes

byte[] data = File.ReadAllBytes("document.hwp");
using var doc = UnhwpDocument.ParseBytes(data);

Tables as CSV

using var doc = UnhwpDocument.ParseFile("document.hwp");

foreach (var table in doc.GetTables())
{
    File.WriteAllText($"s{table.Section}-t{table.Index}.csv", table.Text);
}

Extract Images

using var doc = UnhwpDocument.ParseFile("document.hwp");

foreach (var id in doc.GetResourceIds())
{
    var data = doc.GetResourceData(id);
    if (data != null)
        File.WriteAllBytes(Path.Combine("output", id), data);
}

Handling Failures

UnhwpException.Kind says why a call failed, so you can react to the reason instead of matching on message text:

try
{
    using var doc = UnhwpDocument.ParseFile(path);
}
catch (UnhwpException ex)
{
    switch (ex.Kind)
    {
        case UnhwpErrorKind.OleContainer:
        case UnhwpErrorKind.ZipArchive:
            Console.Error.WriteLine("The file is damaged.");
            break;
        case UnhwpErrorKind.UnknownFormat:
        case UnhwpErrorKind.UnsupportedFormat:
            Console.Error.WriteLine("Not a supported HWP/HWPX document.");
            break;
        default:
            // Also the right branch for a reason this build has no name for.
            Console.Error.WriteLine($"Extraction failed ({ex.Kind}): {ex.Message}");
            break;
    }
}

The numbers behind UnhwpErrorKind are a stable ABI contract: a new reason takes the next free number and existing ones are never renumbered. Always keep a default branch so an unrecognised value degrades to a generic failure rather than going unhandled. Kind is Other for failures raised by the wrapper itself, and never None (which means success).

Features

  • Fast: Native Rust library
  • Complete: Extracts text, tables, images, and document structure
  • Format Support: HWP 5.0, HWPX, and HWP 3.x (legacy)
  • Trim and Native AOT compatible: no reflection-based serialization

API Reference

UnhwpDocument Class

Implements IDisposable; dispose it to release the native document.

Static Members
  • ParseFile(string path) - Parse a document from a file path
  • ParseBytes(byte[] data) - Parse a document from bytes
  • Version - Library version
Instance Methods
  • ToMarkdown(MarkdownOptions? options = null) - Convert to Markdown
  • ToText() - Convert to plain text
  • ToJson(bool compact = false) - Convert to JSON
  • PlainText() - Get plain text (fast extraction)
  • GetTables(bool tsv = false) - Every table as CSV (RFC 4180), or tab-separated, in reading order (IReadOnlyList<TableText>): Section, Index (its place in the section, from 1) and Text. A merged cell's text is in its top-left position and the positions it covers are empty, so every record has the same number of fields.
  • GetResourceIds() - List of resource IDs
  • GetResourceInfo(string id) - Resource metadata as JsonDocument, or null when absent
  • GetResourceData(string id) - Resource binary data, or null when absent
Properties
  • Title - Document title, or null
  • Author - Document author, or null
  • SectionCount - Number of sections
  • ResourceCount - Number of resources

MarkdownOptions Class

  • IncludeFrontmatter - Include YAML frontmatter with document metadata (default false)
  • EscapeSpecialChars - Escape special Markdown characters (default true)
  • ParagraphSpacing - Add extra spacing between paragraphs (default false)
  • Refine - Apply the lossless Markdown shape-refinement pass after rendering (default false)

Platform Support

  • Windows (x64)
  • Linux (x64, glibc and musl)
  • macOS (x64, ARM64)

Target Frameworks

  • .NET 10.0

License

MIT License - see LICENSE for details.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net10.0

    • No dependencies.

NuGet packages (3)

Showing the top 3 NuGet packages that depend on Unhwp:

Package Downloads
FileFlux

Complete document processing SDK optimized for RAG systems. Transform PDF, DOCX, Excel, PowerPoint, Markdown and other formats into high-quality chunks with intelligent semantic boundary detection. Includes advanced chunking strategies, metadata extraction, and performance optimization.

FileFlux.Core

Pure document extraction SDK for RAG systems. Zero AI dependencies. Extract text from PDF, DOCX, Excel, PowerPoint, Markdown, HTML, and text files. Provides IDocumentReader interface and implementations. Use FileFlux.Core for extraction-only scenarios. For AI-enhanced extraction (image OCR, captioning), use the FileFlux package.

FileFlux.Providers.LMSupply

LMSupply local ONNX model provider for FileFlux: document analysis (summarization, metadata extraction), embeddings, OCR, image captioning, and speech transcription — no API key required.

GitHub repositories (1)

Showing the top 1 popular GitHub repositories that depend on Unhwp:

Repository Stars
KnifeLemon/Filee
Free, open-source file converter. Convert 180+ formats offline by drag & drop — images, PDF, Word, Excel, PowerPoint, HWP/HWPX, video, audio, e-books and archives. No upload, no Office needed.
Version Downloads Last Updated
0.18.0 0 10/10/2026
0.17.0 35 10/10/2026
0.16.0 52 10/9/2026
0.15.2 58 10/9/2026
0.15.1 78 10/8/2026
0.15.0 410 10/7/2026
0.14.0 383 10/7/2026
0.13.3 605 10/6/2026
0.13.2 140 10/4/2026
0.13.1 1,683 10/3/2026
0.13.0 3,308 9/28/2026
0.12.0 2,243 9/22/2026
0.11.0 2,971 9/11/2026
0.10.0 132 9/9/2026
0.9.1 138 8/21/2026
0.9.0 127 8/20/2026
0.8.1 122 8/20/2026
0.8.0 162 7/31/2026
0.7.0 134 7/30/2026
0.6.0 138 7/22/2026
Loading failed