FileFlux 0.33.1
See the version list below for details.
dotnet add package FileFlux --version 0.33.1
NuGet\Install-Package FileFlux -Version 0.33.1
<PackageReference Include="FileFlux" Version="0.33.1" />
<PackageVersion Include="FileFlux" Version="0.33.1" />
<PackageReference Include="FileFlux" />
paket add FileFlux --version 0.33.1
#r "nuget: FileFlux, 0.33.1"
#:package FileFlux@0.33.1
#addin nuget:?package=FileFlux&version=0.33.1
#tool nuget:?package=FileFlux&version=0.33.1
FileFlux
.NET document processing library for RAG systems
Overview
FileFlux is a .NET library that transforms various document formats into optimized chunks for RAG (Retrieval-Augmented Generation) systems. Built on high-performance Rust FFI libraries for document parsing.
Key Features
- 5-Stage Stateful Pipeline: Extract → Rule-Refine → LLM-Refine → Chunk → Enrich
- Native Document Readers: Rust FFI-based readers (Unpdf, Undoc, Unhwp) for 2-5x faster processing. Binaries are NuGet-pinned for reproducibility; runtime self-update from GitHub releases is opt-in (off by default — set
UndocNativeLoader.AutoUpdateEnabled = true/UnhwpNativeLoader.AutoUpdateEnabled = trueor theFILEFLUX_NATIVE_AUTOUPDATE=1environment variable) - Multiple Document Formats: PDF, DOCX, XLSX, PPTX, HWP, HWPX, Markdown, HTML, TXT, JSON, CSV
- Chunking Strategies: Auto, Sentence, Paragraph, Token, Hierarchical, Semantic (
ChunkingStrategies— see Chunking Strategies) - Interface-Driven AI: Define AI service interfaces, implement with your preferred provider
- Document Graph: Inter-chunk relationship tracking with sequential, hierarchical, and semantic edges
- Structural Metadata: HeadingPath, page numbers, ContextDependency scores for enhanced RAG
- Language Detection: Automatic language detection using NTextCat
- Document Metadata Extraction (standalone service, not a pipeline stage):
AIMetadataEnricher—EnrichAsync(content, MetadataSchema)returns topics, keywords, a description and schema-specific fields (General, ProductManual, TechnicalDoc) through yourIDocumentAnalysisService, with a rule-based fallback and a content cache. Construct it with aRuleBasedMetadataExtractorand anIMemoryCacheand call it on the text you want described - IEnrichedChunk Interface: Standardized interface for RAG system integration
- Extensible Architecture: Interface-based design for easy customization
- Async Processing: Streaming and parallel processing for large documents
Installation
Full RAG Pipeline
dotnet add package FileFlux
Extraction Only (Minimal Dependencies)
dotnet add package FileFlux.Core
Package Comparison:
| Feature | FileFlux.Core | FileFlux |
|---------|---------------|----------|
| Document Readers (PDF, DOCX, etc.) | ✅ | ✅ |
| Core Interfaces & Models | ✅ | ✅ |
| AI Service Interfaces (IDocumentAnalysisService, IImageToTextService, IAudioToTextService, IEmbeddingService) | ❌ | ✅ |
| Chunking Strategies | ❌ | ✅ |
| FluxCurator & FluxImprover | ❌ | ✅ |
| DocumentProcessor | ❌ | ✅ |
| Use Case | Custom chunking | Full RAG pipeline |
FileFlux.Providers.LMSupply adds local AI implementations of those interfaces (below), and
FileFlux.CLI is a command-line tool over the same pipeline (dotnet tool install -g FileFlux.CLI).
Quick Start
Basic Usage
AddFileFlux() registers an IDocumentProcessorFactory. A processor handles one document: create it from a path (or a
Stream/byte[] with its extension), run the pipeline, and read the chunks from Result — which is itself the chunk
sequence.
using FileFlux;
using FileFlux.Core;
using Microsoft.Extensions.DependencyInjection;
var services = new ServiceCollection();
services.AddFileFlux(); // no logger or AI service required
using var provider = services.BuildServiceProvider();
var factory = provider.GetRequiredService<IDocumentProcessorFactory>();
await using var processor = factory.Create("document.pdf");
await processor.ProcessAsync(); // Extract → Refine → (LLM refine, skipped without an AI service) → Chunk
foreach (var chunk in processor.Result)
{
Console.WriteLine($"Chunk {chunk.ChunkIndex}: {chunk.Content}");
}
Clean chunk content. Before a chunk is surfaced, FileFlux removes all HTML comments (
) from `chunk.Content`. This covers the internal structural markers FileFlux emits and consumes during boundary detection (,,, etc.), so consumers no longer need their own marker-removal step. Note that any HTML comment authored in your source document is also stripped from chunk content. When a chunk begins with a heading marker, its level (1-6) is preserved inchunk.Props[ChunkPropsKeys.HierarchyHeadingLevel].Markdown link reference definitions (
[label]: https://…) are likewise excluded from chunk content — they are document metadata, not body text. The referencing link keeps its display text and resolved target inline, so no content is lost.Known limitation: if a chunker splits a marker across a chunk boundary (e.g.
), neither half matches and both leak — identical to a downstreamregex.
Streaming Processing
ProcessStreamAsync yields chunks as they are produced; Result.Chunks holds all of them once the enumeration ends.
var factory = provider.GetRequiredService<IDocumentProcessorFactory>();
await using var processor = factory.Create("document.pdf");
await foreach (var chunk in processor.ProcessStreamAsync())
{
Console.WriteLine($"Chunk {chunk.ChunkIndex}: {chunk.Content.Length} chars");
}
Chunking Options
var factory = provider.GetRequiredService<IDocumentProcessorFactory>();
await using var processor = factory.Create("document.pdf");
await processor.ProcessAsync(new ProcessingOptions
{
Chunking = new ChunkingOptions
{
Strategy = ChunkingStrategies.Auto, // see Chunking Strategies below
MaxChunkSize = 512, // maximum chunk size (default 1024)
OverlapSize = 64 // overlap between chunks (default 128)
}
});
Stateful Pipeline (v0.9.0+)
The new stateful pipeline provides explicit control over each processing stage:
using FileFlux.Core;
// Create processor via factory
var factory = provider.GetRequiredService<IDocumentProcessorFactory>();
using var processor = factory.Create("document.pdf");
// Execute stages explicitly
await processor.ExtractAsync(); // Stage 1: Raw content extraction
await processor.RefineAsync(); // Stage 2: Rule-based text cleaning
await processor.LlmRefineAsync(); // Stage 3: LLM-powered refinement (optional)
await processor.ChunkAsync(); // Stage 4: Content chunking
await processor.EnrichAsync(); // Stage 5: LLM-powered enrichment (optional)
// Access results at each stage
Console.WriteLine($"State: {processor.State}");
Console.WriteLine($"Raw text length: {processor.Result.Raw?.Text.Length}");
Console.WriteLine($"Sections found: {processor.Result.Refined?.Sections.Count}");
Console.WriteLine($"Chunks created: {processor.Result.Chunks?.Count}");
// Or run full pipeline at once
await processor.ProcessAsync(new ProcessingOptions
{
IncludeEnrich = true,
Enrich = new EnrichOptions { BuildGraph = true }
});
// Text you extracted earlier and stored: start at the Extracted stage, keep the source locations
using var stored = factory.Create(new RawContent
{
Text = storedText,
Spans = storedSpans, // SourceSpan(start, end) { Page = 3 } / { StartTime = …, EndTime = … }
File = new SourceFileInfo { Name = "report.md", Extension = ".md" },
});
await stored.ChunkAsync(); // each chunk's Location.StartPage/EndPage/StartTime/EndTime comes from the spans
// Access the document graph
if (processor.Result.Graph != null)
{
Console.WriteLine($"Graph nodes: {processor.Result.Graph.NodeCount}");
Console.WriteLine($"Graph edges: {processor.Result.Graph.EdgeCount}");
}
Pipeline Stages:
| Stage | Interface | AI | Description |
|-------|-----------|:--:|-------------|
| Extract | IDocumentReader | ❌ | Raw content extraction from files |
| Rule-Refine | IDocumentRefiner | ❌ | Text cleaning, normalization, structure analysis |
| LLM-Refine | ILlmRefiner | ✅ | AI-powered noise removal, sentence restoration |
| Chunk | IChunkerFactory | Optional | Content segmentation with various strategies |
| Enrich | IDocumentEnricher | ✅ | LLM-powered summaries, keywords, contextual text |
AI Service Interfaces
FileFlux defines AI service interfaces - consumer applications provide implementations.
Available Interfaces
| Interface | Purpose | Example Implementations |
|---|---|---|
IDocumentAnalysisService |
Text generation, intelligent chunking. Override GenerateAsync(prompt, GenerationSettings, ct) so LlmRefineOptions/ParsingOptions Temperature/MaxTokens reach your model, and throw GenerationTruncatedException (or, from a service on the shared completion port, its base Flux.Abstractions.TextCompletionTruncatedException) when the model stops at the token limit so a cut-off rewrite is never adopted. Set ProviderInfo.MaxContextLength and the refiner will not send a pass that cannot fit |
OpenAI, Anthropic, LMSupply |
IImageToTextService |
Image captioning, OCR | OpenAI Vision, LMSupply Captioner/OCR |
IAudioToTextService |
Speech transcription — makes audio files readable (0.30.0+); without one, audio is unsupported | LMSupply Transcriber |
IEmbeddingService |
Embedding generation for your own use. The pipeline does not consume it yet — Semantic chunking takes FluxCurator's IEmbedder (see Chunking Strategies) |
OpenAI, LMSupply Embedder |
OpenAICompatibleDocumentAnalysisService (in FileFlux) is a ready IDocumentAnalysisService for OpenAI, Azure OpenAI,
Ollama and other OpenAI-compatible endpoints.
Example: Custom AI Provider
using FileFlux;
using Microsoft.Extensions.DependencyInjection;
var services = new ServiceCollection();
// Your implementations, registered before or after AddFileFlux() — the pipeline resolves them when it runs
services.AddSingleton<IDocumentAnalysisService>(myAnalysisService); // LLM refine, enrichment
services.AddSingleton<IImageToTextService>(myVisionService); // images inside documents
services.AddFileFlux();
AddFileFlux(ServiceLifetime.Singleton) registers the pipeline with the lifetime of a singleton or hosted consumer (default
Scoped); AddDocumentReader<T>() / AddDocumentParser<T>() add your own reader or parser for a format — an added reader wins
over the built-in one for its extensions, registered before or after AddFileFlux() (AddNativeOfficeReader() uses this for the
native DOCX/XLSX/PPTX reader).
Local AI with LMSupply (v0.20.0+)
For local AI processing without external API calls or keys, reference the
FileFlux.Providers.LMSupply package (built on
LMSupply) — it mirrors the
FluxIndex.Providers.LMSupply convention used elsewhere in the ecosystem:
dotnet add package FileFlux.Providers.LMSupply
using FileFlux;
using FileFlux.Providers.LMSupply.Extensions;
using Microsoft.Extensions.DependencyInjection;
var services = new ServiceCollection();
services.AddLMSupplyDocumentAnalysis(); // "default": LMSupply picks a GGUF model for this host
services.AddLMSupplyEmbedding("default"); // an IEmbeddingService for your own use (not read by the pipeline)
services.AddLMSupplyCaptioner(); // or AddLMSupplyOcr() for scanned/text-bearing images
services.AddLMSupplyTranscriber(); // .wav/.mp3 become readable; chunks carry Location.StartTime/EndTime
// services.AddLMSupplyTranscriber(configure: o => { o.Diarize = true; o.NumSpeakers = 3; }); // label who said what
services.AddFileFlux();
An ONNX Runtime GenAI model id (for example microsoft/Phi-4-mini-instruct-onnx) also needs the
LMSupply.Generator.Onnx package, registered with OnnxGeneratorBackend.Register() at startup.
LMSupplyServiceFactory (also in this package) offers lazy/cached service creation with
download-progress reporting for interactive apps like the FileFlux CLI, which consumes this
package the same way rather than carrying its own copy.
Note: FileFlux.Providers.LMSupply is a separate, optional package — FileFlux itself has no
LMSupply dependency. Fully custom providers (Example above) remain the way to integrate any other
AI backend.
Supported Document Formats
| Format | Extension | Reader | Features |
|---|---|---|---|
| Unpdf (Rust FFI) | Text, tables, image extraction | ||
| Word | .docx | Undoc (Rust FFI) | Style and structure preservation |
| Excel | .xlsx | Undoc (Rust FFI) | Multi-sheet and table structure |
| Excel (legacy) | .xls | Built-in (ExcelDataReader) | BIFF binary workbooks; per-sheet markdown tables; CP949 (EUC-KR) fallback for codepage-less BIFF5/7 |
Mislabelled workbooks (since 0.17.0) — the two Excel readers route on the container's magic bytes rather than the declared extension, in both directions: a compound-file (
.xls) workbook named.xlsxextracts through the legacy reader, and an OOXML package named.xlsextracts through the OOXML one.RawContent.File.Extensionreports the container that was actually parsed, not the name the file arrived under. Content that is neither container fails withextraction_failure_reason=container_mismatchinstead of the ZIP parser's "could not find EOCD", which reads as corruption when the file is simply not a workbook. | PowerPoint | .pptx | Undoc (Rust FFI) | Slide and notes extraction | | HWP | .hwp, .hwpx | Unhwp (Rust FFI) | Native Korean document support | | Markdown | .md | Built-in | Structure preservation | | HTML | .html, .htm | Built-in | Web content extraction | | CSV/TSV | .csv, .tsv | Built-in (CsvHelper) | Header-aware markdown table serialization; UTF-8/BOM + CP949 (EUC-KR) fallback decoding | | Text | .txt, .json | Built-in | Basic text processing | | Audio | .wav, .mp3 |IAudioToTextService(e.g.AddLMSupplyTranscriber()) | Speech as text, one paragraph per segment (speaker-labelled when the service separates speakers); chunks carryLocation.StartTime/EndTime. Unsupported when no service is registered |
Known Limitations
PDF Processing
- Vector Graphics Tables: Tables created with drawing primitives (lines/rectangles) instead of text layout may not be detected. These are rendered as images in most PDF viewers.
- Complex Multi-column Layouts: Documents with intricate multi-column arrangements may have suboptimal text ordering.
- Scanned Documents: OCR is not included; scanned PDFs require pre-processing with external OCR tools. When a PDF parses but yields no text at all, the reader returns empty content with
Hints["extraction_failure_reason"]set to"no_text_layer"(image-only/scanned — pages draw images without a readable text layer, via Unpdf page introspection) or"blank_page"(no text or image content at all), plus an explanatory warning — so consumers can classify these distinctly from parse errors. - Partial Extraction: When whole-document extraction fails, FileFlux automatically falls back to per-page extraction. Pages that cannot be extracted are skipped and recorded in
RawContent.Errors.RawContent.Statusis set toProcessingStatus.Partialwhen some pages fail, allowing RAG pipelines to use the successfully extracted content rather than losing the entire document. - Parse Failures:
extraction_failure_reasonabove covers documents that parse but yield no text;extraction_error_kindcovers documents that fail to parse, naming Unpdf's structured failure classification (PdfParse,UnknownFormat,Encrypted,Corrupted,Io,MissingObject, …) so consumers can classify without matching on message prose. On partial extraction it arrives asHints["extraction_error_kind"], listing every distinct kind seen across the skipped pages ("PdfParse+MissingObject"). When extraction fails entirely there is noRawContentto carry hints, so the same value is embedded in the thrownDocumentProcessingException.Messageas aextraction_error_kind=<kind>token:
try
{
var content = await reader.ExtractAsync("scan.pdf");
if (content.Hints.TryGetValue("extraction_error_kind", out var kind))
logger.LogWarning("Some pages unreadable: {Kind}", kind); // e.g. "PdfParse"
}
catch (DocumentProcessingException ex)
{
// ex.Message contains "... [extraction_error_kind=Corrupted]"
logger.LogError(ex, "PDF could not be parsed");
}
A value from a newer native build passes through as its number rather than being dropped, so unknown kinds stay reportable. Kinds numbered 100 and above are raised at the native library's interop boundary rather than by the document, so a failure carrying one is a library-side problem to report upstream, not a defect in the file — the thrown message says so instead of filing it under parse failure.
- Incomplete Extraction: A damaged PDF does not always fail. If part of its page tree cannot be read, the parser recovers the rest and extraction succeeds over a shorter document. FileFlux flags that with
Hints["pages_incomplete"] = trueandRawContent.Status = ProcessingStatus.Partial, plus a warning — so a page that never arrived is not indexed as a page that never existed.Hints["declared_page_count"]carries the count the document declares, to compare against the extractedpage_count. The flag is deliberately a boolean and never a loss figure: one unresolved page-tree node can cost a single page or a whole subtree, so the number of lost pages is not knowable.
var content = await reader.ExtractAsync("damaged.pdf");
if (content.Hints.ContainsKey("pages_incomplete"))
{
// declared_page_count is absent when the document's own declaration was unreadable —
// itself a damage signal, so the flag can be set without a number to compare against.
content.Hints.TryGetValue("declared_page_count", out var declared);
logger.LogWarning("Indexing an incomplete document: declared {Declared}, extracted {Extracted}",
declared ?? "unknown", content.Hints["page_count"]);
}
ReadAsync (stage 0) carries the same signal in DocumentProps, with ReadResult.Status set to Partial — that stage reports the page count, so it is where a short page set most easily passes for a whole document.
- Suppressed Text Runs: Some PDFs use fonts whose character codes the decoder cannot resolve — emitting the raw bytes would produce mojibake rather than text, so the decoder discards the run instead. FileFlux surfaces that with
Hints["suppressed_text_runs"](the discarded run count) andRawContent.Status = ProcessingStatus.Partial, plus a warning. When every run in the document was discarded, the empty result getsHints["extraction_failure_reason"] = "text_runs_suppressed"instead of"no_text_layer"— this is not a scanned document, and does not need OCR.
var content = await reader.ExtractAsync("broken-font.pdf");
if (content.Hints.TryGetValue("suppressed_text_runs", out var runs))
logger.LogWarning("Document lost {Runs} text run(s) to an unresolvable font", runs);
ReadAsync (stage 0) carries the same signal in DocumentProps, with ReadResult.Status set to Partial.
Table Extraction
FileFlux uses layout-based table detection with confidence scoring:
- Tables with confidence score ≥ 0.5 are converted to Markdown format
- Low-confidence tables fall back to plain text to prevent garbled output
- Table quality metrics are exposed via
RawContent.Hintsfor consumer applications
Document-Specific Notes
- Excel: Very large worksheets (>100K rows) may impact memory usage
- PowerPoint: Embedded objects are extracted as placeholder text
- HTML: JavaScript-rendered content is not supported
Chunking Strategies
| Strategy | Output characteristics | Prerequisites |
|---|---|---|
Auto (default) |
Resolved to a concrete strategy by content analysis: short text → Sentence, 4+ paragraphs → Paragraph, sentence-structured → Sentence, otherwise Token. The resolved strategy is logged and recorded in each chunk's Strategy. |
— |
Sentence |
Sentence-boundary chunks, language-aware | — |
Paragraph |
Paragraph-boundary chunks; best for Markdown/blogs; oversized paragraphs fall back to sentence splits | — |
Token |
Token-budget chunks for unstructured text | — |
Hierarchical |
Heading-structure-aware chunks | — |
Semantic |
Embedding-similarity boundaries | Requires a FluxCurator IEmbedder in the container — otherwise chunker creation throws ArgumentException |
Structural metadata: every ProcessAsync/ChunkAsync chunk carries Location.StartChar/EndChar
(offsets into the refined text), Location.HeadingPath/Section (hierarchical heading context,
e.g. Root Title > Sub Section), and Props[ChunkPropsKeys.HierarchyPath] ("hierarchy.path") — on the streaming ChunkStreamAsync too.
Location.StartPage/EndPage name the first and last source page of the chunk's text for PDFs (0.30.0+), and
Location.StartTime/EndTime carry the source time range for timed sources.
These come from RawContent.Spans: a reader describes where each stretch of its text came from (SourceSpan —
Page, StartTime/EndTime), refinement carries the spans onto RefinedContent.Spans, and chunking writes the
spans each chunk overlaps onto its Location. A custom IDocumentReader fills Spans to get the same. Spans are
dropped (and chunk pages left null) when a step rebuilds the text from scratch — table/block conversion from
structured reader output, or an LLM rewrite.
Advanced Features
AI services are optional — see AI Service Interfaces. With an IDocumentAnalysisService registered,
the LLM refine stage runs (noise removal, sentence restoration) and EnrichAsync builds summaries, keywords and the document
graph; without one those stages are skipped. 📖 See Tutorial for AI service implementation examples.
📊 Quality Analysis
Evaluate and optimize chunking quality for RAG systems:
using FileFlux.Infrastructure.Quality;
// Score the chunks a processing run produced
var factory = provider.GetRequiredService<IDocumentProcessorFactory>();
await using var processor = factory.Create("document.pdf");
await processor.ProcessAsync();
var metrics = await ChunkQualityEngine.CalculateQualityMetricsAsync(processor.Result);
Console.WriteLine($"Completeness: {metrics.AverageCompleteness:P0}");
Console.WriteLine($"Boundaries: {metrics.BoundaryQuality:P0}");
Console.WriteLine($"Size spread: {metrics.SizeDistribution:P0}");
To compare strategies, run ProcessAsync once per ChunkingOptions.Strategy (a new processor each time) and compare the metrics.
📖 See Architecture for quality analysis details.
Documentation
- Tutorial - Detailed usage guide and examples
- Architecture - System design and pipeline documentation
- Changelog - Version history and release notes
Project Structure
FileFlux/
├── src/
│ ├── FileFlux.Core/ # Extraction only (zero AI dependencies)
│ │ ├── Contracts/ # IDocumentProcessor, ProcessingResult
│ │ ├── Core/ # IDocumentRefiner, IDocumentEnricher
│ │ └── Domain/ # DocumentGraph, RefinedContent, StructuredElement
│ ├── FileFlux/ # Full RAG pipeline (interface-driven)
│ │ └── Infrastructure/ # StatefulDocumentProcessor, DocumentRefiner, DocumentEnricher
│ └── FileFlux.Providers.LMSupply/ # Optional local-AI provider package (v0.20.0+)
├── cli/ # CLI (published: `dotnet tool install -g FileFlux.CLI`)
│ └── FileFlux.CLI/
├── tests/
│ └── FileFlux.Tests/ # Test suite (343+ tests)
└── samples/
└── FileFlux.SampleApp/ # Usage examples
Contributing
- Create and discuss an issue
- Work on a feature branch
- Add/modify tests
- Submit a pull request
License
MIT License - See LICENSE file
Support
- Issue Reports: GitHub Issues
- Feature Requests: GitHub Discussions
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- CsvHelper (>= 33.1.0)
- ExcelDataReader (>= 3.7.0)
- FileFlux.Core (>= 0.33.1)
- Flux.Abstractions (>= 0.26.0)
- FluxCurator (>= 0.9.1)
- FluxImprover (>= 0.14.14)
- HtmlAgilityPack (>= 1.12.4)
- Markdig (>= 1.3.2)
- Microsoft.Extensions.Caching.Memory (>= 10.0.12)
- Microsoft.Extensions.Configuration (>= 10.0.12)
- Microsoft.Extensions.DependencyInjection.Abstractions (>= 10.0.12)
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.12)
- Microsoft.Extensions.ObjectPool (>= 10.0.12)
- Microsoft.Extensions.Options.ConfigurationExtensions (>= 10.0.12)
- NTextCat (>= 0.3.65)
- Undoc (>= 0.13.0)
- Unhwp (>= 0.13.0)
- Unpdf (>= 0.22.0)
NuGet packages (5)
Showing the top 5 NuGet packages that depend on FileFlux:
| Package | Downloads |
|---|---|
|
FluxFeed
FluxFeed - document pipeline surface (ingest/parse/clean) feeding FluxIndex. File-source vault with git-like tracking and real-time folder monitoring. |
|
|
FluxIndex.Integrations.FileFlux
FileFlux document parsing/chunking integration for FluxIndex — DI wiring and the document processing pipeline. |
|
|
IronHive.Flux.Core
Core adapters bridging IronHive AI services to Flux ecosystem (FileFlux, WebFlux, FluxIndex) |
|
|
FluxIndex.Extensions.FileVault
FluxIndex FileVault - Git-like file tracking system for RAG indexing with real-time folder monitoring |
|
|
FileFlux.Providers.LMSupply
LMSupply local ONNX model provider for FileFlux: document analysis (summarization, metadata extraction), embeddings, OCR, image captioning, and speech transcription — no API key required. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.33.3 | 0 | 9/29/2026 |
| 0.33.2 | 0 | 9/29/2026 |
| 0.33.1 | 0 | 9/28/2026 |
| 0.33.0 | 0 | 9/28/2026 |
| 0.32.0 | 0 | 9/28/2026 |
| 0.31.13 | 0 | 9/28/2026 |
| 0.31.12 | 73 | 9/28/2026 |
| 0.31.11 | 66 | 9/27/2026 |
| 0.31.10 | 52 | 9/27/2026 |
| 0.31.9 | 41 | 9/27/2026 |
| 0.31.8 | 92 | 9/27/2026 |
| 0.31.7 | 96 | 9/27/2026 |
| 0.31.6 | 59 | 9/27/2026 |
| 0.31.5 | 69 | 9/27/2026 |
| 0.31.4 | 89 | 9/26/2026 |
| 0.31.3 | 104 | 9/26/2026 |
| 0.31.2 | 52 | 9/26/2026 |
| 0.31.1 | 88 | 9/26/2026 |
| 0.31.0 | 76 | 9/26/2026 |
| 0.30.0 | 79 | 9/26/2026 |
Complete FileFlux SDK with document readers, intelligent chunking strategies, and RAG optimization - Now in a single unified package