Independent Submission K. Rickert Internet-Draft PipeStream AI Intended status: Informational 23 July 2026 Expires: 24 January 2027 PipeStream Application Profile for Distributed Document Processing draft-krickert-pipestream-docproc-00 Abstract This document defines an Application Profile for the PipeStream core protocol, mapping its generic recursive scatter-gather semantics to the domain of distributed document processing, AI ingestion, and retrieval-augmented generation (RAG) pipelines. This profile defines the concrete semantics for the four PipeStream Data Layers: BlobBag (Layer 0), SemanticLayer (Layer 1), ParsedData (Layer 2), and CustomEntity (Layer 3). It also specifies the PipeDoc application-level envelope, ownership contexts for multi-tenant environments, and profile conventions for document-oriented processing and archival correlation. About This Document This note is to be removed before publishing as an RFC. Status information for this document may be found at https://datatracker.ietf.org/doc/draft-krickert-pipestream-docproc/. Discussion of this document takes place on the Individual Group mailing list (mailto:kristian.rickert@pipestream.ai). Status of This Memo This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79. Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet- Drafts is at https://datatracker.ietf.org/drafts/current/. Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress." Rickert Expires 24 January 2027 [Page 1] Internet-Draft PipeStream DocProc July 2026 This Internet-Draft will expire on 24 January 2027. Copyright Notice Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved. This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/ license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Table of Contents 1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . 3 1.1. Purpose . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.2. Relationship to PipeStream Core . . . . . . . . . . . . . 3 1.3. Terminology . . . . . . . . . . . . . . . . . . . . . . . 3 2. PipeDoc Entity Envelope . . . . . . . . . . . . . . . . . . . 4 2.1. Core Fields . . . . . . . . . . . . . . . . . . . . . . . 4 2.2. OwnershipContext . . . . . . . . . . . . . . . . . . . . 5 3. Data Layer Definitions . . . . . . . . . . . . . . . . . . . 6 3.1. Layer 0: BlobBag . . . . . . . . . . . . . . . . . . . . 7 3.2. Layer 1: SemanticLayer . . . . . . . . . . . . . . . . . 7 3.3. Layer 2: ParsedData . . . . . . . . . . . . . . . . . . . 7 3.4. Layer 3: CustomEntity . . . . . . . . . . . . . . . . . . 8 4. Document Processing Pipeline Conventions . . . . . . . . . . 8 4.1. PARSE Stage (Document Decomposition) . . . . . . . . . . 8 4.2. PROCESS Stage . . . . . . . . . . . . . . . . . . . . . . 8 4.3. SINK Stage . . . . . . . . . . . . . . . . . . . . . . . 8 5. Security Considerations . . . . . . . . . . . . . . . . . . . 8 5.1. PII in Document Metadata . . . . . . . . . . . . . . . . 9 5.2. Multi-Tenancy via OwnershipContext . . . . . . . . . . . 9 5.3. Malicious Content and Resource Exhaustion . . . . . . . . 9 5.4. Metadata and Structured Content Injection . . . . . . . . 9 5.5. Cross-Tenant Trust Boundaries . . . . . . . . . . . . . . 9 6. IANA Considerations . . . . . . . . . . . . . . . . . . . . . 10 6.1. Profile Identification . . . . . . . . . . . . . . . . . 10 7. Normative References . . . . . . . . . . . . . . . . . . . . 10 Appendix A. Complete CDDL Schema . . . . . . . . . . . . . . . . 10 Appendix B. Example Processing Patterns . . . . . . . . . . . . 13 B.1. Text Extraction . . . . . . . . . . . . . . . . . . . . . 13 B.2. NLP Enrichment . . . . . . . . . . . . . . . . . . . . . 13 B.3. Structured Table Extraction . . . . . . . . . . . . . . . 13 B.4. Image Processing . . . . . . . . . . . . . . . . . . . . 13 B.5. Example Sink Patterns . . . . . . . . . . . . . . . . . . 13 Author's Address . . . . . . . . . . . . . . . . . . . . . . . . 13 Rickert Expires 24 January 2027 [Page 2] Internet-Draft PipeStream DocProc July 2026 1. Introduction 1.1. Purpose This document defines an Application Profile for PipeStream [PIPESTREAM] that specifies entity payload formats and processing semantics for distributed document processing pipelines. This profile assigns concrete meanings to PipeStream's four data layers, defines the PipeDoc application-level entity envelope, and specifies interoperable payload conventions for document ingestion, enrichment, and indexing workflows. 1.2. Relationship to PipeStream Core PipeStream Core defines the transport mapping, control stream framing, recursive entity lifecycle, and resilience semantics. This profile does not modify any PipeStream Core wire format. Instead, it defines how document-processing implementations interpret the payload bytes carried within PipeStream entities. This document is intended as an independent industry profile rather than as a standards-track extension to PipeStream Core. It can evolve on a faster cadence than the core transport specification while preserving wire compatibility with PipeStream entities and control frames. Implementations of this profile MUST implement PipeStream Core [PIPESTREAM] Layer 0 at minimum. Implementations that require recursive document decomposition, such as processing embedded documents within archives, SHOULD implement Layer 1. Implementations that interact with external services, such as third-party NLP APIs or human review workflows, SHOULD implement Layer 2. 1.3. Terminology The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here. This profile uses all capitalized PipeStream Core terms as defined in [PIPESTREAM]. In addition, the following profile terms are used: *PipeDoc* The application-level document envelope carried within an entity payload. PipeDoc provides a stable document identifier and ownership context for document-processing pipelines. Rickert Expires 24 January 2027 [Page 3] Internet-Draft PipeStream DocProc July 2026 *Profile Version* A profile-level schema version carried within PipeDoc. It identifies which revision of this document defined the payload layout for a given document-processing entity. *BlobBag* The Layer 0 representation for raw binary document content and related attachments. *SemanticLayer* The Layer 1 representation for annotated content, chunking output, embeddings, and NLP annotations. *ParsedData* The Layer 2 representation for structured extracted metadata, including fields and tables. *CustomEntity* The Layer 3 representation for application-specific extensions to this profile. 2. PipeDoc Entity Envelope 2.1. Core Fields This profile defines PipeDoc as the application-level envelope carried within document-processing entity payloads. Rickert Expires 24 January 2027 [Page 4] Internet-Draft PipeStream DocProc July 2026 +===============+=================+===========+=====================+ |Field |Type |Requirement|Description | +===============+=================+===========+=====================+ |profile-version|uint |REQUIRED |Version of this | | | | |document-processing | | | | |profile used to | | | | |encode the envelope | +---------------+-----------------+-----------+---------------------+ |doc-id |tstr |REQUIRED |Stable document | | | | |identifier for the | | | | |pipeline run or | | | | |deduplicated source | | | | |object | +---------------+-----------------+-----------+---------------------+ |entity-id |uint |REQUIRED |Profile-visible copy | | | | |of the enclosing | | | | |PipeStream Entity ID | | | | |for archival and | | | | |off-transport | | | | |correlation | +---------------+-----------------+-----------+---------------------+ |ownership |ownership-context|OPTIONAL |Multi-tenant | | | | |ownership and access | | | | |metadata | +---------------+-----------------+-----------+---------------------+ Table 1 Field names and types follow the CDDL schema in Appendix A. The entity-id field in PipeDoc MUST match the entity-id value in the enclosing PipeStream Entity Header. This field is retained so that archived payloads, detached artifacts, and reprocessed document representations can be correlated after transport headers have been stripped or normalized away. PipeDoc MAY carry additional document metadata and layer-specific payload structures as defined in Appendix A. The profile-version field MUST be set to 1 for payloads conforming to this document. 2.2. OwnershipContext OwnershipContext provides application-layer multi-tenancy and authorization metadata for document-processing deployments. It is not interpreted by PipeStream Core; it is consumed only by implementations of this profile and related applications. Rickert Expires 24 January 2027 [Page 5] Internet-Draft PipeStream DocProc July 2026 +===========+=======+=============+================================+ | Field | Type | Requirement | Description | +===========+=======+=============+================================+ | tenant-id | tstr | OPTIONAL | Administrative tenant or | | | | | account boundary | +-----------+-------+-------------+--------------------------------+ | owner-id | tstr | OPTIONAL | Individual owner or service | | | | | principal | +-----------+-------+-------------+--------------------------------+ | acl | [* | OPTIONAL | Access control principals | | | tstr] | | allowed to access the document | +-----------+-------+-------------+--------------------------------+ Table 2 Implementations MAY omit OwnershipContext in single-tenant or otherwise trusted environments. 3. Data Layer Definitions The PipeStream Entity Header layer field is authoritative for the semantic interpretation of a document-processing payload. Every document-processing entity carries a PipeDoc envelope, but exactly one layer-specific payload family is expected to be populated for a given entity: +=======+==========================+=======================+ | Layer | PipeDoc fields expected | Notes | +=======+==========================+=======================+ | 0 | blob_bag plus shared | Raw source content | | | envelope metadata | | +-------+--------------------------+-----------------------+ | 1 | semantic_result plus | Annotated or enriched | | | shared envelope metadata | intermediate output | +-------+--------------------------+-----------------------+ | 2 | parsed-metadata and/or | Structured extraction | | | structured-data plus | results | | | shared envelope metadata | | +-------+--------------------------+-----------------------+ | 3 | custom_entity plus | Profile extension or | | | shared envelope metadata | vendor-specific | | | | payload | +-------+--------------------------+-----------------------+ Table 3 Rickert Expires 24 January 2027 [Page 6] Internet-Draft PipeStream DocProc July 2026 An implementation of this profile SHOULD NOT populate multiple layer- specific payload families in the same PipeDoc instance. Shared envelope metadata such as profile-version, doc-id, entity-id, search- metadata, and ownership MAY appear at any layer. 3.1. Layer 0: BlobBag BlobBag is the Layer 0 representation for raw binary document data. It holds one or more blobs that together represent the original source material entering the pipeline, such as PDFs, images, office attachments, or archive members. Each Blob MAY embed bytes inline or MAY reference externally stored data via the FileStorageReference type defined in PipeStream Core. Blob metadata MAY include MIME type, filename, size, and checksum information. 3.2. Layer 1: SemanticLayer SemanticLayer is the Layer 1 representation for annotated content produced by semantic processing stages. Typical contents include: * chunked text segments * dense vector embeddings * named-entity annotations * model metadata and chunking strategy metadata This layer is intended for enrichment stages that transform raw bytes into semantically meaningful segments while preserving enough context for downstream search, retrieval, and ranking systems. 3.3. Layer 2: ParsedData ParsedData is the Layer 2 representation for structured extraction output. Typical contents include: * extracted key-value fields * normalized metadata attributes * extracted tables * parser-specific raw textual output Rickert Expires 24 January 2027 [Page 7] Internet-Draft PipeStream DocProc July 2026 This layer is intended for downstream systems that require normalized records rather than raw documents or semantic chunks. 3.4. Layer 3: CustomEntity CustomEntity is the Layer 3 representation for application-specific extensions. This profile reserves Layer 3 for payloads that build on the document-processing model but require data structures not standardized by this document. Receivers that implement this profile but do not understand a Layer 3 payload MAY pass it through unchanged, provided that PipeStream Core processing requirements are still met. 4. Document Processing Pipeline Conventions This section describes common conventions used by document-processing deployments that implement this profile. It is intentionally non- prescriptive: implementations MAY realize these stages using different internal service boundaries, model stacks, and sink targets. 4.1. PARSE Stage (Document Decomposition) In this profile, the PARSE stage commonly performs document decomposition. A root document entity MAY be split into child entities representing pages, embedded documents, archive members, images, or other logical subcomponents. The decomposition strategy is implementation-specific but SHOULD preserve enough lineage metadata to allow rehydration at later stages. 4.2. PROCESS Stage The PROCESS stage transforms document entities between profile- defined layer representations. Typical examples include text extraction, OCR, semantic chunking, embedding generation, entity recognition, and structured field or table extraction. Example processing patterns are described in Appendix B. 4.3. SINK Stage The SINK stage represents terminal consumption of document-processing results. Common sink patterns include indexing, archival storage, and workflow notification, but this profile does not require a fixed sink registry or specific backend products. 5. Security Considerations Rickert Expires 24 January 2027 [Page 8] Internet-Draft PipeStream DocProc July 2026 5.1. PII in Document Metadata This profile commonly carries document metadata that may contain personally identifiable information or commercially sensitive content. Implementations SHOULD minimize exposure of titles, filenames, semantic annotations, and extracted fields to only those processing stages that require them. When documents are routed through multi-tenant infrastructure, implementations SHOULD encrypt sensitive application-level metadata at rest and SHOULD avoid exposing unnecessary identifiers in operational logs. 5.2. Multi-Tenancy via OwnershipContext OwnershipContext is an application-layer construct and therefore is not protected by PipeStream Core semantics beyond QUIC transport security. Implementations that rely on OwnershipContext for authorization MUST treat it as security-sensitive metadata and validate it against local policy before granting access to document payloads or derived results. 5.3. Malicious Content and Resource Exhaustion Document-processing pipelines routinely ingest untrusted payloads. Implementations SHOULD defend against malformed archives, zip bombs, polyglot files, decompression attacks, path traversal in embedded filenames, and parser-specific exploit inputs. BlobBag processing SHOULD apply bounded resource policies for maximum object size, expansion ratios, page counts, recursion depth, and extraction time. 5.4. Metadata and Structured Content Injection Titles, filenames, annotations, extracted fields, and table content MAY contain attacker-controlled strings. Implementations SHOULD treat these values as untrusted input when rendering search results, constructing queries, invoking downstream tools, or generating prompts for language models. 5.5. Cross-Tenant Trust Boundaries Shared indexing, storage, or enrichment infrastructure can create cross-tenant leakage risks if document-processing metadata is reused outside its intended scope. Implementations SHOULD isolate tenant data paths, avoid sharing authorization context across tenants, and ensure that cached semantic or structured outputs are keyed by both doc-id and the relevant ownership boundary. Rickert Expires 24 January 2027 [Page 9] Internet-Draft PipeStream DocProc July 2026 6. IANA Considerations This document requests no IANA actions. 6.1. Profile Identification This profile is identified by the case-sensitive string DOCPROC. Implementations MAY advertise this identifier in out-of-band configuration, capability metadata, or application-specific routing tables. The initial profile schema version defined by this document is 1. Receivers that do not support the advertised profile-version in a PipeDoc payload SHOULD reject that payload at the application layer rather than attempting best-effort interpretation. 7. Normative References [PIPESTREAM] Rickert, K., "PipeStream: A Recursive Entity Streaming Protocol for Distributed Processing over QUIC", Work in Progress, Internet-Draft, draft-krickert-pipestream-03, July 2026, . [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March 1997, . [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017, . [RFC8610] Birkholz, H., Vigano, C., and C. Bormann, "Concise Data Definition Language (CDDL): A Notational Convention to Express Concise Binary Object Representation (CBOR) and JSON Data Structures", RFC 8610, DOI 10.17487/RFC8610, June 2019, . Appendix A. Complete CDDL Schema This appendix provides the profile's consolidated CDDL [RFC8610] definitions. The file-storage-reference and encryption-metadata types are defined by reference in PipeStream Core Appendix C and are reused here without modification. Rickert Expires 24 January 2027 [Page 10] Internet-Draft PipeStream DocProc July 2026 pipe-doc = { profile-version: uint, doc-id: tstr, entity-id: uint, ? search-metadata: search-metadata, ? blob-bag: blob-bag, ? semantic-result: semantic-processing-result, ? layer2-payload: layer2-payload, ? custom-entity: any, ? ownership: ownership-context, } search-metadata = { ? title: tstr, ? keywords: [* tstr], ? description: tstr, ? custom-fields: { * tstr => tstr }, } blob-bag = blob / blobs blobs = { blobs: [* blob], } blob = { blob-id: tstr, ? drive-id: tstr, content: bstr / file-storage-reference, ? mime-type: tstr, ? filename: tstr, ? size-bytes: int, ? checksum: tstr, ? checksum-type: checksum-type, } checksum-type = &( unspecified: 0, md5: 1, sha1: 2, sha256: 3, sha512: 4, ) semantic-processing-result = { ? chunks: [* semantic-chunk], ? chunking-strategy: tstr, ? processing-metadata: { * tstr => tstr }, Rickert Expires 24 January 2027 [Page 11] Internet-Draft PipeStream DocProc July 2026 } semantic-chunk = { chunk-id: tstr, ? chunk-number: int, ? embedding-info: chunk-embedding, ? metadata: { * tstr => any }, ? annotations: [* nlp-annotation], } chunk-embedding = { text-content: tstr, ? vector: [* float], ? model-id: tstr, ? original-char-start-offset: int, ? original-char-end-offset: int, } nlp-annotation = { type: tstr, label: tstr, ? start-offset: int, ? end-offset: int, ? confidence: float, ? attributes: { * tstr => tstr }, } parsed-metadata = { parser-id: tstr, ? fields: { * tstr => any }, ? tables: [* table-data], ? raw-output: tstr, } layer2-payload = { ? parsed-metadata: { * tstr => parsed-metadata }, ? structured-data: any, } table-data = { table-id: tstr, ? headers: [* tstr], ? rows: [* table-row], } table-row = { cells: [* tstr], } Rickert Expires 24 January 2027 [Page 12] Internet-Draft PipeStream DocProc July 2026 ownership-context = { ? tenant-id: tstr, ? owner-id: tstr, ? acl: [* tstr], } Appendix B. Example Processing Patterns This appendix is non-normative. B.1. Text Extraction A text-extraction stage transforms binary document content into textual or layout-aware intermediate representations. Typical outputs include page text, OCR results, or format-specific structural markup. B.2. NLP Enrichment An enrichment stage adds semantic metadata such as chunking, embeddings, named entities, classifications, or relation annotations. These results are commonly encoded in SemanticLayer payloads. B.3. Structured Table Extraction A table-extraction stage identifies structured tabular regions and emits normalized table representations suitable for indexing or analytics. B.4. Image Processing An image-processing stage derives metadata or features from document images, such as OCR overlays, captions, detections, or classification results. B.5. Example Sink Patterns Common sink patterns include search indexing, archival persistence, and workflow notification. Backend-specific integrations such as particular search engines, object stores, or message buses are deployment choices, not protocol requirements of this profile. Author's Address Kristian Rickert PipeStream AI Email: kristian.rickert@pipestream.ai Rickert Expires 24 January 2027 [Page 13]