WebFlux 0.17.0

There is a newer version of this package available.
See the version list below for details.
dotnet add package WebFlux --version 0.17.0
                    
NuGet\Install-Package WebFlux -Version 0.17.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="WebFlux" Version="0.17.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="WebFlux" Version="0.17.0" />
                    
Directory.Packages.props
<PackageReference Include="WebFlux" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add WebFlux --version 0.17.0
                    
#r "nuget: WebFlux, 0.17.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package WebFlux@0.17.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=WebFlux&version=0.17.0
                    
Install as a Cake Addin
#tool nuget:?package=WebFlux&version=0.17.0
                    
Install as a Cake Tool

WebFlux

A .NET SDK for preprocessing web content for RAG (Retrieval-Augmented Generation) systems.

NuGet Version NuGet Downloads .NET Support License

Overview

WebFlux processes web content into chunks optimized for RAG systems. It handles web crawling, content extraction and chunking.

Installation

dotnet add package WebFlux
Package What it is
WebFlux Crawling, extraction and chunking — browser-free. AddWebFlux()
WebFlux.Playwright Playwright-backed rendering for pages that only render with JavaScript. AddWebFluxPlaywright() after AddWebFlux(), then CrawlOptions.UseDynamicRendering = true

Quick Start

AddWebFlux() needs no AI service: crawling, extraction and chunking run as they are.

using Microsoft.Extensions.DependencyInjection;
using WebFlux.Core.Interfaces;
using WebFlux.Core.Options;
using WebFlux.Extensions;

var services = new ServiceCollection();
services.AddWebFlux();

await using var provider = services.BuildServiceProvider();
var processor = provider.GetRequiredService<IWebContentProcessor>();

// One page
var chunks = await processor.ProcessUrlAsync("https://example.com");
foreach (var chunk in chunks)
{
    Console.WriteLine($"Chunk {chunk.SequenceNumber}: {chunk.Content}");
}

// A site: pass CrawlOptions — without them ProcessWebsiteAsync processes the start page only
var crawl = new CrawlOptions { MaxDepth = 2, MaxPages = 20 };
await foreach (var chunk in processor.ProcessWebsiteAsync("https://example.com", crawl))
{
    Console.WriteLine($"{chunk.SourceUrl} #{chunk.SequenceNumber}: {chunk.Content}");
}

Features

  • Crawling — ProcessWebsiteAsync(url, CrawlOptions) streams chunks as pages arrive. CrawlOptions.Strategy: BreadthFirst (default), DepthFirst, Sitemap (reads sitemap.xml), Dynamic (needs WebFlux.Playwright); UseDynamicRendering = true routes to Dynamic whatever the strategy. MaxDepth / MaxPages bound the crawl.
  • Chunking strategies — Auto, Smart, Semantic, Paragraph, FixedSize, MemoryOptimized (ChunkingOptions.Strategy, see below); an unknown strategy name throws and lists the available ones.
  • Content formats — HTML, Markdown, JSON, XML and plain text.
  • Web standards — robots.txt, matched per RFC 9309 — groups, longest-match with Allow winning ties, * and $. Honoured on every crawl entry point and on the extract API; turn it off per call with CrawlOptions.RespectRobotsTxt or ExtractOptions.RespectRobotsTxt (both default to on). A refusal is reported as CrawlResult.DisallowedByRobotsTxt / ExtractErrorCodes.DisallowedByRobotsTxt rather than as a failed request.
  • Identity — every request carries one User-Agent — CrawlOptions.UserAgent, default WebFluxUserAgent.Default (WebFlux/{version} (+https://github.com/iyulab/WebFlux)) — and the robots.txt group is chosen by its product token, so "MyBot/1.0" follows User-agent: MyBot. CrawlOptions.CustomHeaders are sent per request.
  • Request timeouts — CrawlOptions.TimeoutMs (default 30 000) and ExtractOptions.TimeoutSeconds (default 15) bound every request a call makes — the page, its robots.txt, a sitemap — on the HTTP and the Playwright crawlers alike. A timeout is not retried, whatever MaxRetries says: it arrives as CrawlResult.TimedOut / ExtractErrorCodes.Timeout after about the time you asked for, from CrawlAsync, CrawlWebsiteAsync and CrawlSitemapAsync alike. A non-positive value is not "no timeout": the call throws ArgumentException naming the option before any request is made.
  • Retries — MaxRetries (on CrawlOptions for a crawl, on ExtractOptions for a single URL) counts attempts once, never nested. Transport errors, 408, 429 and 5xx are retried with backoff; a 404/403/410, a robots.txt refusal and a timeout are final.
  • Rich metadata — SEO, Open Graph, Schema.org and Twitter Cards.
    • On the crawl path (ProcessWebsiteAsync): CrawlOptions.UseHtmlMetadata (default on, no AI) attaches the HTML snapshot to Metadata.HtmlMetadata; CrawlOptions.EnableMetadataExtraction (opt-in) runs the AI extractor with MetadataSchema/CustomMetadataPrompt when an ITextCompletionService (or your own IWebMetadataExtractor) is registered, sampling long pages to MetadataExtractionMaxChars (title + headings + first N characters).
  • Events — IEventPublisher (see below) for processing, per-URL and extraction events.
  • Configuration — AddWebFlux(config => …) or AddWebFlux(configuration) (the "WebFlux" section) sets the defaults ProcessUrlAsync starts from: crawling, chunking, AI enhancement.

Chunking Strategies

Strategy Use Case
Auto Automatically selects best strategy based on content
Smart Structured HTML documentation
Semantic General web pages and articles — needs a FluxCurator IEmbedder registered before AddWebFlux()
MemoryOptimized Same token-based splitting as FixedSize; kept as a name for large-document callers
Paragraph Markdown with natural boundaries
FixedSize Uniform chunks for testing

Chunking runs on FluxCurator. For Semantic, register its embedder (FluxCurator.Core.Core.IEmbedder) in the container before AddWebFlux(); the chunker factory picks it up.

Services You Can Provide

WebFlux uses the Interface Provider pattern: nothing below is required for crawling, extraction and chunking.

ITextCompletionService (optional)

LLM text completion, from the shared Flux.Abstractions contract package — only CompleteAsync is required, the rest have default implementations. With one registered:

  • AI metadata on the crawl path — CrawlOptions.EnableMetadataExtraction = true (above).
  • AI enhancement (summaries, rewrites) on ProcessUrlAsync — register the enhancement service and turn it on in the configuration:
using Flux.Abstractions;
using Microsoft.Extensions.DependencyInjection;
using WebFlux.Extensions;

services.AddScoped<ITextCompletionService, MyCompletionService>();
services.AddWebFlux(config => config.AiEnhancement.Enabled = true);
services.AddWebFluxAIEnhancement();

Implement WebFlux.Core.Interfaces.IWebLlmService instead (it extends ITextCompletionService with IsAvailableAsync and GetHealthInfo) when you also want health checks; AddWebFluxAIServices<TService>() registers it as the completion service and adds the enhancement service.

Main Processor

IWebContentProcessor is the entry point; it is also registered as the two focused interfaces below.

using WebFlux.Core.Options;

// Single URL
var chunks = await processor.ProcessUrlAsync("https://example.com");

// Website crawling (streaming)
await foreach (var chunk in processor.ProcessWebsiteAsync(url, crawlOptions, chunkOptions))
{
    // Process chunk
}

// Several URLs: chunks per URL
var results = await processor.ProcessUrlsBatchAsync(urls, chunkOptions);

// HTML you already have
var fromHtml = await processor.ProcessHtmlAsync("<html>…</html>", "https://example.com/page", chunkOptions);

For consumers that only need extraction or chunking:

// Extraction only (single URL, batch, or streamed batch: ExtractBatchAsync / ExtractBatchStreamAsync)
var extractor = provider.GetRequiredService<IContentExtractService>();
var result = await extractor.ExtractContentAsync("https://example.com");

// Chunking only
var chunker = provider.GetRequiredService<IContentChunkService>();
var chunks = await chunker.ProcessUrlAsync("https://example.com");

Extensibility

IChunkingStrategy

Implement WebFlux.Core.Interfaces.IChunkingStrategy — Name, Description and ChunkAsync(ExtractedContent, ChunkingOptions?, CancellationToken) returning IReadOnlyList<WebContentChunk>.

IEventPublisher

Subscribe to pipeline events for monitoring and metrics. IEventPublisher is registered as a singleton by AddWebFlux().

using WebFlux.Core.Interfaces;
using WebFlux.Core.Models.Events;

var publisher = provider.GetRequiredService<IEventPublisher>();

using var started = publisher.Subscribe<UrlProcessingStartedEvent>(e => Console.WriteLine($"Processing {e.Url}"));
using var failed = publisher.Subscribe<ContentExtractionFailedEvent>(e => Console.WriteLine($"{e.Url}: {e.Error}"));
using var done = publisher.Subscribe<ProcessingCompletedEvent>(e =>
    Console.WriteLine($"{e.ProcessedChunkCount} chunks in {e.TotalProcessingTime}"));

// Or every event
using var all = publisher.SubscribeAll(e =>
{
    logger.LogInformation("WebFlux event {EventType}", e.EventType);
    return Task.CompletedTask;
});

Published events (WebFlux.Core.Models.Events):

When Events
ProcessUrlAsync / the configured pipeline ProcessingStartedEvent, ProcessingProgressEvent, ProcessingCompletedEvent
Every URL a crawler fetches UrlProcessingStartedEvent
Every extraction ContentExtractionStartedEvent, ContentExtractionCompletedEvent, ContentExtractionFailedEvent
Retries and circuit breaking ResilienceEvent

The namespace declares more event types (crawling, chunking and monitoring); they are not published yet. All events derive from ProcessingEvent (EventId, EventType, Timestamp, Severity, CorrelationId).

Configuration

using WebFlux.Core.Options;

var options = new CrawlOptions
{
    MaxDepth = 3,
    MaxPages = 100,
    RespectRobotsTxt = true,
    UserAgent = "MyBot/1.0",
    TimeoutMs = 10_000          // per request; a timeout is not retried
};

var chunkOptions = new ChunkingOptions
{
    Strategy = ChunkingStrategyType.Auto,
    MaxChunkSize = 512,
    ChunkOverlap = 64
};

await foreach (var chunk in processor.ProcessWebsiteAsync(url, options, chunkOptions))
{
    // Handle chunk
}

Documentation

License

MIT License - see LICENSE file for details.

Support

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (5)

Showing the top 5 NuGet packages that depend on WebFlux:

Package Downloads
IronHive.DeepResearch

Deep Research module - autonomous research agent system

IronHive.Flux.Core

Core adapters bridging IronHive AI services to Flux ecosystem (FileFlux, WebFlux, FluxIndex)

FluxIndex.Integrations.WebFlux

WebFlux web-content ingestion integration for FluxIndex — DI wiring and context builder extensions.

IronHive.Flux.WebLookup

WebLookup → WebFlux → FluxIndex RAG pipeline for discovering, processing, and indexing web content

WebFlux.Playwright

Playwright-backed dynamic rendering for WebFlux. Add this package only when JavaScript-rendered pages must be crawled; WebFlux itself stays browser-free.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.19.5 203 10/2/2026
0.19.4 145 10/1/2026
0.19.3 437 9/30/2026
0.19.2 122 9/30/2026
0.19.1 192 9/29/2026
0.19.0 145 9/29/2026
0.18.0 199 9/29/2026
0.17.0 222 9/29/2026
0.16.0 2,460 9/24/2026
0.15.0 258 9/23/2026
0.14.1 251 9/23/2026
0.14.0 405 9/22/2026
0.13.0 183 9/22/2026
0.12.0 284 9/21/2026
0.11.0 164 9/21/2026
0.10.0 185 9/21/2026
0.9.0 244 9/20/2026
0.7.4 535 9/18/2026
0.7.3 263 9/17/2026
Loading failed