WebFlux 0.19.1
See the version list below for details.
dotnet add package WebFlux --version 0.19.1
NuGet\Install-Package WebFlux -Version 0.19.1
<PackageReference Include="WebFlux" Version="0.19.1" />
<PackageVersion Include="WebFlux" Version="0.19.1" />
<PackageReference Include="WebFlux" />
paket add WebFlux --version 0.19.1
#r "nuget: WebFlux, 0.19.1"
#:package WebFlux@0.19.1
#addin nuget:?package=WebFlux&version=0.19.1
#tool nuget:?package=WebFlux&version=0.19.1
WebFlux
A .NET SDK for preprocessing web content for RAG (Retrieval-Augmented Generation) systems.
Overview
WebFlux processes web content into chunks optimized for RAG systems. It handles web crawling, content extraction and chunking.
Installation
dotnet add package WebFlux
| Package | What it is |
|---|---|
WebFlux |
Crawling, extraction and chunking — browser-free. AddWebFlux() |
WebFlux.Playwright |
Playwright-backed rendering for pages that only render with JavaScript. AddWebFluxPlaywright() after AddWebFlux(), then CrawlOptions.UseDynamicRendering = true |
Quick Start
AddWebFlux() needs no AI service: crawling, extraction and chunking run as they are.
using Microsoft.Extensions.DependencyInjection;
using WebFlux.Core.Interfaces;
using WebFlux.Core.Options;
using WebFlux.Extensions;
var services = new ServiceCollection();
services.AddWebFlux();
await using var provider = services.BuildServiceProvider();
var processor = provider.GetRequiredService<IWebContentProcessor>();
// One page
var chunks = await processor.ProcessUrlAsync("https://example.com");
foreach (var chunk in chunks)
{
Console.WriteLine($"Chunk {chunk.SequenceNumber}: {chunk.Content}");
}
// A site: pass CrawlOptions — without them ProcessWebsiteAsync processes the start page only
var crawl = new CrawlOptions { MaxDepth = 2, MaxPages = 20 };
await foreach (var chunk in processor.ProcessWebsiteAsync("https://example.com", crawl))
{
Console.WriteLine($"{chunk.SourceUrl} #{chunk.SequenceNumber}: {chunk.Content}");
}
Features
- Crawling —
ProcessWebsiteAsync(url, CrawlOptions)streams chunks as pages arrive.CrawlOptions.Strategy:BreadthFirst(default),DepthFirst,Sitemap(reads sitemap.xml),Dynamic(needsWebFlux.Playwright);UseDynamicRendering = trueroutes toDynamicwhatever the strategy.MaxDepth/MaxPagesbound the crawl. - Chunking strategies — Auto, Smart, Semantic, Paragraph, FixedSize, MemoryOptimized (
ChunkingOptions.Strategy, see below); an unknown strategy name throws and lists the available ones. - Content formats — HTML, Markdown, JSON, XML and plain text.
- Web standards — robots.txt, matched per RFC 9309
— groups, longest-match with
Allowwinning ties,*and$. Honoured on every crawl entry point and on the extract API; turn it off per call withCrawlOptions.RespectRobotsTxtorExtractOptions.RespectRobotsTxt(both default to on). A refusal is reported asCrawlResult.DisallowedByRobotsTxt/ExtractErrorCodes.DisallowedByRobotsTxtrather than as a failed request. - Identity — every request carries one User-Agent —
CrawlOptions.UserAgent, defaultWebFluxUserAgent.Default(WebFlux/{version} (+https://github.com/iyulab/WebFlux)) — and the robots.txt group is chosen by its product token, so"MyBot/1.0"followsUser-agent: MyBot.CrawlOptions.CustomHeadersare sent per request. - Request timeouts —
CrawlOptions.TimeoutMs(default 30 000) andExtractOptions.TimeoutSeconds(default 15) bound every request a call makes — the page, its robots.txt, a sitemap — on the HTTP and the Playwright crawlers alike. A timeout is not retried, whateverMaxRetriessays: it arrives asCrawlResult.TimedOut/ExtractErrorCodes.Timeoutafter about the time you asked for, fromCrawlAsync,CrawlWebsiteAsyncandCrawlSitemapAsyncalike. A non-positive value is not "no timeout": the call throwsArgumentExceptionnaming the option before any request is made. - Retries —
MaxRetries(onCrawlOptionsfor a crawl, onExtractOptionsfor a single URL) counts attempts once, never nested. Transport errors, 408, 429 and 5xx are retried with backoff; a 404/403/410, a robots.txt refusal and a timeout are final. - Rich metadata — SEO, Open Graph, Schema.org and Twitter Cards.
- On the crawl path (
ProcessWebsiteAsync):CrawlOptions.UseHtmlMetadata(default on, no AI) attaches the HTML snapshot toMetadata.HtmlMetadata;CrawlOptions.EnableMetadataExtraction(opt-in) runs the AI extractor withMetadataSchema/CustomMetadataPromptwhen anITextCompletionService(or your ownIWebMetadataExtractor) is registered, sampling long pages toMetadataExtractionMaxChars(title + headings + first N characters).
- On the crawl path (
- Events —
IEventPublisher(see below) for processing, per-chunk, per-URL and extraction events. - Configuration —
AddWebFlux(config => …)orAddWebFlux(configuration)(the"WebFlux"section) sets the defaultsProcessUrlAsyncstarts from: crawling, chunking, AI enhancement.
Chunking Strategies
| Strategy | Use Case |
|---|---|
| Auto | Automatically selects best strategy based on content |
| Smart | Structured HTML documentation |
| Semantic | General web pages and articles — needs a FluxCurator IEmbedder registered before AddWebFlux() |
| MemoryOptimized | Same token-based splitting as FixedSize; kept as a name for large-document callers |
| Paragraph | Markdown with natural boundaries |
| FixedSize | Uniform chunks for testing |
Chunking runs on FluxCurator. For Semantic, register its embedder
(FluxCurator.Core.Core.IEmbedder) in the container before AddWebFlux(); the chunker factory picks it up.
Services You Can Provide
WebFlux uses the Interface Provider pattern: nothing below is required for crawling, extraction and chunking.
ITextCompletionService (optional)
LLM text completion, from the shared Flux.Abstractions
contract package — only CompleteAsync is required, the rest have default implementations. With one registered:
- AI metadata on the crawl path —
CrawlOptions.EnableMetadataExtraction = true(above). - AI enhancement (summaries, rewrites) on
ProcessUrlAsync— register the enhancement service and turn it on in the configuration:
using Flux.Abstractions;
using Microsoft.Extensions.DependencyInjection;
using WebFlux.Extensions;
services.AddScoped<ITextCompletionService, MyCompletionService>();
services.AddWebFlux(config => config.AiEnhancement.Enabled = true);
services.AddWebFluxAIEnhancement();
Implement WebFlux.Core.Interfaces.IWebLlmService instead (it extends ITextCompletionService with
IsAvailableAsync and GetHealthInfo) when you also want health checks; AddWebFluxAIServices<TService>() registers
it as the completion service and adds the enhancement service.
Main Processor
IWebContentProcessor is the entry point; it is also registered as the two focused interfaces below.
using WebFlux.Core.Options;
// Single URL
var chunks = await processor.ProcessUrlAsync("https://example.com");
// Website crawling (streaming)
await foreach (var chunk in processor.ProcessWebsiteAsync(url, crawlOptions, chunkOptions))
{
// Process chunk
}
// Several URLs: chunks per URL
var results = await processor.ProcessUrlsBatchAsync(urls, chunkOptions);
// HTML you already have
var fromHtml = await processor.ProcessHtmlAsync("<html>…</html>", "https://example.com/page", chunkOptions);
For consumers that only need extraction or chunking:
// Extraction only (single URL, batch, or streamed batch: ExtractBatchAsync / ExtractBatchStreamAsync)
var extractor = provider.GetRequiredService<IContentExtractService>();
var result = await extractor.ExtractContentAsync("https://example.com");
// Chunking only
var chunker = provider.GetRequiredService<IContentChunkService>();
var chunks = await chunker.ProcessUrlAsync("https://example.com");
Extensibility
IChunkingStrategy
Implement WebFlux.Core.Interfaces.IChunkingStrategy — Name, Description and
ChunkAsync(ExtractedContent, ChunkingOptions?, CancellationToken) returning IReadOnlyList<WebContentChunk>.
IEventPublisher
Subscribe to pipeline events for monitoring and metrics. IEventPublisher is registered as a singleton by AddWebFlux().
using WebFlux.Core.Interfaces;
using WebFlux.Core.Models.Events;
var publisher = provider.GetRequiredService<IEventPublisher>();
using var started = publisher.Subscribe<UrlProcessingStartedEvent>(e => Console.WriteLine($"Processing {e.Url}"));
using var failed = publisher.Subscribe<ContentExtractionFailedEvent>(e => Console.WriteLine($"{e.Url}: {e.Error}"));
using var done = publisher.Subscribe<ProcessingCompletedEvent>(e =>
Console.WriteLine($"{e.ProcessedChunkCount} chunks in {e.TotalProcessingTime}"));
// Or every event
using var all = publisher.SubscribeAll(e =>
{
logger.LogInformation("WebFlux event {EventType}", e.EventType);
return Task.CompletedTask;
});
Published events (WebFlux.Core.Models.Events):
| When | Events |
|---|---|
Every processing run (ProcessUrlAsync, ProcessWebsiteAsync, the configured pipeline) |
ProcessingStartedEvent, then ChunkGeneratedEvent per chunk and ProcessingProgressEvent every 10 chunks, then ProcessingCompletedEvent — or ProcessingFailedEvent when the run throws |
| Every URL a crawler fetches | UrlProcessingStartedEvent, then UrlProcessedEvent (success status) or UrlProcessingFailedEvent (error status, or no response) |
| Every extraction | ContentExtractionStartedEvent, ContentExtractionCompletedEvent, ContentExtractionFailedEvent |
| Retries and circuit breaking | ResilienceEvent |
A subscription to a base type receives every event derived from it — SubscribeAll and Subscribe<ProcessingEvent> see them all.
A subscriber that throws is counted in GetStatistics().PublishErrors and does not interrupt the run or the other subscribers.
All events derive from ProcessingEvent (EventId, EventType, Timestamp, Severity, CorrelationId).
Configuration
using WebFlux.Core.Options;
var options = new CrawlOptions
{
MaxDepth = 3,
MaxPages = 100,
RespectRobotsTxt = true,
UserAgent = "MyBot/1.0",
TimeoutMs = 10_000 // per request; a timeout is not retried
};
var chunkOptions = new ChunkingOptions
{
Strategy = ChunkingStrategyType.Auto,
MaxChunkSize = 512,
ChunkOverlap = 64
};
await foreach (var chunk in processor.ProcessWebsiteAsync(url, options, chunkOptions))
{
// Handle chunk
}
Documentation
- Tutorial - Step-by-step guide with practical examples
- Architecture - System design and pipeline
- Interfaces - API contracts and implementation guide
- Chunking Strategies - Detailed strategy guide
- Changelog - Version history and release notes
License
MIT License - see LICENSE file for details.
Support
- Issues: GitHub Issues
- Package: NuGet
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- AngleSharp (>= 1.5.2)
- Flux.Abstractions (>= 0.26.0)
- FluxCurator (>= 0.10.0)
- FluxCurator.Core (>= 0.10.0)
- HtmlAgilityPack (>= 1.12.4)
- Markdig (>= 1.2.0)
- Microsoft.Extensions.Caching.Abstractions (>= 10.0.12)
- Microsoft.Extensions.Caching.Memory (>= 10.0.12)
- Microsoft.Extensions.Configuration (>= 10.0.12)
- Microsoft.Extensions.Configuration.Abstractions (>= 10.0.12)
- Microsoft.Extensions.Configuration.Binder (>= 10.0.12)
- Microsoft.Extensions.DependencyInjection (>= 10.0.12)
- Microsoft.Extensions.DependencyInjection.Abstractions (>= 10.0.12)
- Microsoft.Extensions.Http (>= 10.0.12)
- Microsoft.Extensions.Logging (>= 10.0.12)
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.12)
- Polly (>= 8.6.6)
- Polly.Extensions.Http (>= 3.0.0)
- ReverseMarkdown (>= 6.2.1)
- YamlDotNet (>= 18.1.0)
NuGet packages (5)
Showing the top 5 NuGet packages that depend on WebFlux:
| Package | Downloads |
|---|---|
|
IronHive.DeepResearch
Deep Research module - autonomous research agent system |
|
|
IronHive.Flux.Core
Core adapters bridging IronHive AI services to Flux ecosystem (FileFlux, WebFlux, FluxIndex) |
|
|
FluxIndex.Integrations.WebFlux
WebFlux web-content ingestion integration for FluxIndex — DI wiring and context builder extensions. |
|
|
IronHive.Flux.WebLookup
WebLookup → WebFlux → FluxIndex RAG pipeline for discovering, processing, and indexing web content |
|
|
WebFlux.Playwright
Playwright-backed dynamic rendering for WebFlux. Add this package only when JavaScript-rendered pages must be crawled; WebFlux itself stays browser-free. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.19.5 | 312 | 10/2/2026 |
| 0.19.4 | 160 | 10/1/2026 |
| 0.19.3 | 473 | 9/30/2026 |
| 0.19.2 | 131 | 9/30/2026 |
| 0.19.1 | 204 | 9/29/2026 |
| 0.19.0 | 150 | 9/29/2026 |
| 0.18.0 | 206 | 9/29/2026 |
| 0.17.0 | 232 | 9/29/2026 |
| 0.16.0 | 2,466 | 9/24/2026 |
| 0.15.0 | 260 | 9/23/2026 |
| 0.14.1 | 253 | 9/23/2026 |
| 0.14.0 | 407 | 9/22/2026 |
| 0.13.0 | 185 | 9/22/2026 |
| 0.12.0 | 286 | 9/21/2026 |
| 0.11.0 | 167 | 9/21/2026 |
| 0.10.0 | 187 | 9/21/2026 |
| 0.9.0 | 246 | 9/20/2026 |
| 0.7.4 | 537 | 9/18/2026 |
| 0.7.3 | 265 | 9/17/2026 |