Skip to main content

PipeStream Application Profile for Distributed Document Processing
draft-krickert-pipestream-docproc-00

Document Type Active Internet-Draft (individual)
Author Kristian E Rickert
Last updated 2026-07-23
RFC stream (None)
Intended RFC status (None)
Formats
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state I-D Exists
Telechat date (None)
Responsible AD (None)
Send notices to (None)
draft-krickert-pipestream-docproc-00
Independent Submission                                        K. Rickert
Internet-Draft                                             PipeStream AI
Intended status: Informational                              23 July 2026
Expires: 24 January 2027

   PipeStream Application Profile for Distributed Document Processing
                  draft-krickert-pipestream-docproc-00

Abstract

   This document defines an Application Profile for the PipeStream core
   protocol, mapping its generic recursive scatter-gather semantics to
   the domain of distributed document processing, AI ingestion, and
   retrieval-augmented generation (RAG) pipelines.

   This profile defines the concrete semantics for the four PipeStream
   Data Layers: BlobBag (Layer 0), SemanticLayer (Layer 1), ParsedData
   (Layer 2), and CustomEntity (Layer 3).  It also specifies the PipeDoc
   application-level envelope, ownership contexts for multi-tenant
   environments, and profile conventions for document-oriented
   processing and archival correlation.

About This Document

   This note is to be removed before publishing as an RFC.

   Status information for this document may be found at
   https://datatracker.ietf.org/doc/draft-krickert-pipestream-docproc/.

   Discussion of this document takes place on the Individual Group
   mailing list (mailto:kristian.rickert@pipestream.ai).

Status of This Memo

   This Internet-Draft is submitted in full conformance with the
   provisions of BCP 78 and BCP 79.

   Internet-Drafts are working documents of the Internet Engineering
   Task Force (IETF).  Note that other groups may also distribute
   working documents as Internet-Drafts.  The list of current Internet-
   Drafts is at https://datatracker.ietf.org/drafts/current/.

   Internet-Drafts are draft documents valid for a maximum of six months
   and may be updated, replaced, or obsoleted by other documents at any
   time.  It is inappropriate to use Internet-Drafts as reference
   material or to cite them other than as "work in progress."

Rickert                  Expires 24 January 2027                [Page 1]
Internet-Draft             PipeStream DocProc                  July 2026

   This Internet-Draft will expire on 24 January 2027.

Copyright Notice

   Copyright (c) 2026 IETF Trust and the persons identified as the
   document authors.  All rights reserved.

   This document is subject to BCP 78 and the IETF Trust's Legal
   Provisions Relating to IETF Documents (https://trustee.ietf.org/
   license-info) in effect on the date of publication of this document.
   Please review these documents carefully, as they describe your rights
   and restrictions with respect to this document.

Table of Contents

   1.  Introduction  . . . . . . . . . . . . . . . . . . . . . . . .   3
     1.1.  Purpose . . . . . . . . . . . . . . . . . . . . . . . . .   3
     1.2.  Relationship to PipeStream Core . . . . . . . . . . . . .   3
     1.3.  Terminology . . . . . . . . . . . . . . . . . . . . . . .   3
   2.  PipeDoc Entity Envelope . . . . . . . . . . . . . . . . . . .   4
     2.1.  Core Fields . . . . . . . . . . . . . . . . . . . . . . .   4
     2.2.  OwnershipContext  . . . . . . . . . . . . . . . . . . . .   5
   3.  Data Layer Definitions  . . . . . . . . . . . . . . . . . . .   6
     3.1.  Layer 0: BlobBag  . . . . . . . . . . . . . . . . . . . .   7
     3.2.  Layer 1: SemanticLayer  . . . . . . . . . . . . . . . . .   7
     3.3.  Layer 2: ParsedData . . . . . . . . . . . . . . . . . . .   7
     3.4.  Layer 3: CustomEntity . . . . . . . . . . . . . . . . . .   8
   4.  Document Processing Pipeline Conventions  . . . . . . . . . .   8
     4.1.  PARSE Stage (Document Decomposition)  . . . . . . . . . .   8
     4.2.  PROCESS Stage . . . . . . . . . . . . . . . . . . . . . .   8
     4.3.  SINK Stage  . . . . . . . . . . . . . . . . . . . . . . .   8
   5.  Security Considerations . . . . . . . . . . . . . . . . . . .   8
     5.1.  PII in Document Metadata  . . . . . . . . . . . . . . . .   9
     5.2.  Multi-Tenancy via OwnershipContext  . . . . . . . . . . .   9
     5.3.  Malicious Content and Resource Exhaustion . . . . . . . .   9
     5.4.  Metadata and Structured Content Injection . . . . . . . .   9
     5.5.  Cross-Tenant Trust Boundaries . . . . . . . . . . . . . .   9
   6.  IANA Considerations . . . . . . . . . . . . . . . . . . . . .  10
     6.1.  Profile Identification  . . . . . . . . . . . . . . . . .  10
   7.  Normative References  . . . . . . . . . . . . . . . . . . . .  10
   Appendix A.  Complete CDDL Schema . . . . . . . . . . . . . . . .  10
   Appendix B.  Example Processing Patterns  . . . . . . . . . . . .  13
     B.1.  Text Extraction . . . . . . . . . . . . . . . . . . . . .  13
     B.2.  NLP Enrichment  . . . . . . . . . . . . . . . . . . . . .  13
     B.3.  Structured Table Extraction . . . . . . . . . . . . . . .  13
     B.4.  Image Processing  . . . . . . . . . . . . . . . . . . . .  13
     B.5.  Example Sink Patterns . . . . . . . . . . . . . . . . . .  13
   Author's Address  . . . . . . . . . . . . . . . . . . . . . . . .  13

Rickert                  Expires 24 January 2027                [Page 2]
Internet-Draft             PipeStream DocProc                  July 2026

1.  Introduction

1.1.  Purpose

   This document defines an Application Profile for PipeStream
   [PIPESTREAM] that specifies entity payload formats and processing
   semantics for distributed document processing pipelines.  This
   profile assigns concrete meanings to PipeStream's four data layers,
   defines the PipeDoc application-level entity envelope, and specifies
   interoperable payload conventions for document ingestion, enrichment,
   and indexing workflows.

1.2.  Relationship to PipeStream Core

   PipeStream Core defines the transport mapping, control stream
   framing, recursive entity lifecycle, and resilience semantics.  This
   profile does not modify any PipeStream Core wire format.  Instead, it
   defines how document-processing implementations interpret the payload
   bytes carried within PipeStream entities.

   This document is intended as an independent industry profile rather
   than as a standards-track extension to PipeStream Core.  It can
   evolve on a faster cadence than the core transport specification
   while preserving wire compatibility with PipeStream entities and
   control frames.

   Implementations of this profile MUST implement PipeStream Core
   [PIPESTREAM] Layer 0 at minimum.  Implementations that require
   recursive document decomposition, such as processing embedded
   documents within archives, SHOULD implement Layer 1.  Implementations
   that interact with external services, such as third-party NLP APIs or
   human review workflows, SHOULD implement Layer 2.

1.3.  Terminology

   The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
   "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
   "OPTIONAL" in this document are to be interpreted as described in BCP
   14 [RFC2119] [RFC8174] when, and only when, they appear in all
   capitals, as shown here.

   This profile uses all capitalized PipeStream Core terms as defined in
   [PIPESTREAM].  In addition, the following profile terms are used:

   *PipeDoc*  The application-level document envelope carried within an
      entity payload.  PipeDoc provides a stable document identifier and
      ownership context for document-processing pipelines.

Rickert                  Expires 24 January 2027                [Page 3]
Internet-Draft             PipeStream DocProc                  July 2026

   *Profile Version*  A profile-level schema version carried within
      PipeDoc.  It identifies which revision of this document defined
      the payload layout for a given document-processing entity.

   *BlobBag*  The Layer 0 representation for raw binary document content
      and related attachments.

   *SemanticLayer*  The Layer 1 representation for annotated content,
      chunking output, embeddings, and NLP annotations.

   *ParsedData*  The Layer 2 representation for structured extracted
      metadata, including fields and tables.

   *CustomEntity*  The Layer 3 representation for application-specific
      extensions to this profile.

2.  PipeDoc Entity Envelope

2.1.  Core Fields

   This profile defines PipeDoc as the application-level envelope
   carried within document-processing entity payloads.

Rickert                  Expires 24 January 2027                [Page 4]
Internet-Draft             PipeStream DocProc                  July 2026

   +===============+=================+===========+=====================+
   |Field          |Type             |Requirement|Description          |
   +===============+=================+===========+=====================+
   |profile-version|uint             |REQUIRED   |Version of this      |
   |               |                 |           |document-processing  |
   |               |                 |           |profile used to      |
   |               |                 |           |encode the envelope  |
   +---------------+-----------------+-----------+---------------------+
   |doc-id         |tstr             |REQUIRED   |Stable document      |
   |               |                 |           |identifier for the   |
   |               |                 |           |pipeline run or      |
   |               |                 |           |deduplicated source  |
   |               |                 |           |object               |
   +---------------+-----------------+-----------+---------------------+
   |entity-id      |uint             |REQUIRED   |Profile-visible copy |
   |               |                 |           |of the enclosing     |
   |               |                 |           |PipeStream Entity ID |
   |               |                 |           |for archival and     |
   |               |                 |           |off-transport        |
   |               |                 |           |correlation          |
   +---------------+-----------------+-----------+---------------------+
   |ownership      |ownership-context|OPTIONAL   |Multi-tenant         |
   |               |                 |           |ownership and access |
   |               |                 |           |metadata             |
   +---------------+-----------------+-----------+---------------------+

                                  Table 1

   Field names and types follow the CDDL schema in Appendix A.

   The entity-id field in PipeDoc MUST match the entity-id value in the
   enclosing PipeStream Entity Header.  This field is retained so that
   archived payloads, detached artifacts, and reprocessed document
   representations can be correlated after transport headers have been
   stripped or normalized away.

   PipeDoc MAY carry additional document metadata and layer-specific
   payload structures as defined in Appendix A.

   The profile-version field MUST be set to 1 for payloads conforming to
   this document.

2.2.  OwnershipContext

   OwnershipContext provides application-layer multi-tenancy and
   authorization metadata for document-processing deployments.  It is
   not interpreted by PipeStream Core; it is consumed only by
   implementations of this profile and related applications.

Rickert                  Expires 24 January 2027                [Page 5]
Internet-Draft             PipeStream DocProc                  July 2026

   +===========+=======+=============+================================+
   | Field     | Type  | Requirement | Description                    |
   +===========+=======+=============+================================+
   | tenant-id | tstr  | OPTIONAL    | Administrative tenant or       |
   |           |       |             | account boundary               |
   +-----------+-------+-------------+--------------------------------+
   | owner-id  | tstr  | OPTIONAL    | Individual owner or service    |
   |           |       |             | principal                      |
   +-----------+-------+-------------+--------------------------------+
   | acl       | [*    | OPTIONAL    | Access control principals      |
   |           | tstr] |             | allowed to access the document |
   +-----------+-------+-------------+--------------------------------+

                                 Table 2

   Implementations MAY omit OwnershipContext in single-tenant or
   otherwise trusted environments.

3.  Data Layer Definitions

   The PipeStream Entity Header layer field is authoritative for the
   semantic interpretation of a document-processing payload.  Every
   document-processing entity carries a PipeDoc envelope, but exactly
   one layer-specific payload family is expected to be populated for a
   given entity:

       +=======+==========================+=======================+
       | Layer | PipeDoc fields expected  | Notes                 |
       +=======+==========================+=======================+
       | 0     | blob_bag plus shared     | Raw source content    |
       |       | envelope metadata        |                       |
       +-------+--------------------------+-----------------------+
       | 1     | semantic_result plus     | Annotated or enriched |
       |       | shared envelope metadata | intermediate output   |
       +-------+--------------------------+-----------------------+
       | 2     | parsed-metadata and/or   | Structured extraction |
       |       | structured-data plus     | results               |
       |       | shared envelope metadata |                       |
       +-------+--------------------------+-----------------------+
       | 3     | custom_entity plus       | Profile extension or  |
       |       | shared envelope metadata | vendor-specific       |
       |       |                          | payload               |
       +-------+--------------------------+-----------------------+

                                 Table 3

Rickert                  Expires 24 January 2027                [Page 6]
Internet-Draft             PipeStream DocProc                  July 2026

   An implementation of this profile SHOULD NOT populate multiple layer-
   specific payload families in the same PipeDoc instance.  Shared
   envelope metadata such as profile-version, doc-id, entity-id, search-
   metadata, and ownership MAY appear at any layer.

3.1.  Layer 0: BlobBag

   BlobBag is the Layer 0 representation for raw binary document data.
   It holds one or more blobs that together represent the original
   source material entering the pipeline, such as PDFs, images, office
   attachments, or archive members.

   Each Blob MAY embed bytes inline or MAY reference externally stored
   data via the FileStorageReference type defined in PipeStream Core.
   Blob metadata MAY include MIME type, filename, size, and checksum
   information.

3.2.  Layer 1: SemanticLayer

   SemanticLayer is the Layer 1 representation for annotated content
   produced by semantic processing stages.  Typical contents include:

   *  chunked text segments

   *  dense vector embeddings

   *  named-entity annotations

   *  model metadata and chunking strategy metadata

   This layer is intended for enrichment stages that transform raw bytes
   into semantically meaningful segments while preserving enough context
   for downstream search, retrieval, and ranking systems.

3.3.  Layer 2: ParsedData

   ParsedData is the Layer 2 representation for structured extraction
   output.  Typical contents include:

   *  extracted key-value fields

   *  normalized metadata attributes

   *  extracted tables

   *  parser-specific raw textual output

Rickert                  Expires 24 January 2027                [Page 7]
Internet-Draft             PipeStream DocProc                  July 2026

   This layer is intended for downstream systems that require normalized
   records rather than raw documents or semantic chunks.

3.4.  Layer 3: CustomEntity

   CustomEntity is the Layer 3 representation for application-specific
   extensions.  This profile reserves Layer 3 for payloads that build on
   the document-processing model but require data structures not
   standardized by this document.

   Receivers that implement this profile but do not understand a Layer 3
   payload MAY pass it through unchanged, provided that PipeStream Core
   processing requirements are still met.

4.  Document Processing Pipeline Conventions

   This section describes common conventions used by document-processing
   deployments that implement this profile.  It is intentionally non-
   prescriptive: implementations MAY realize these stages using
   different internal service boundaries, model stacks, and sink
   targets.

4.1.  PARSE Stage (Document Decomposition)

   In this profile, the PARSE stage commonly performs document
   decomposition.  A root document entity MAY be split into child
   entities representing pages, embedded documents, archive members,
   images, or other logical subcomponents.  The decomposition strategy
   is implementation-specific but SHOULD preserve enough lineage
   metadata to allow rehydration at later stages.

4.2.  PROCESS Stage

   The PROCESS stage transforms document entities between profile-
   defined layer representations.  Typical examples include text
   extraction, OCR, semantic chunking, embedding generation, entity
   recognition, and structured field or table extraction.  Example
   processing patterns are described in Appendix B.

4.3.  SINK Stage

   The SINK stage represents terminal consumption of document-processing
   results.  Common sink patterns include indexing, archival storage,
   and workflow notification, but this profile does not require a fixed
   sink registry or specific backend products.

5.  Security Considerations

Rickert                  Expires 24 January 2027                [Page 8]
Internet-Draft             PipeStream DocProc                  July 2026

5.1.  PII in Document Metadata

   This profile commonly carries document metadata that may contain
   personally identifiable information or commercially sensitive
   content.  Implementations SHOULD minimize exposure of titles,
   filenames, semantic annotations, and extracted fields to only those
   processing stages that require them.

   When documents are routed through multi-tenant infrastructure,
   implementations SHOULD encrypt sensitive application-level metadata
   at rest and SHOULD avoid exposing unnecessary identifiers in
   operational logs.

5.2.  Multi-Tenancy via OwnershipContext

   OwnershipContext is an application-layer construct and therefore is
   not protected by PipeStream Core semantics beyond QUIC transport
   security.  Implementations that rely on OwnershipContext for
   authorization MUST treat it as security-sensitive metadata and
   validate it against local policy before granting access to document
   payloads or derived results.

5.3.  Malicious Content and Resource Exhaustion

   Document-processing pipelines routinely ingest untrusted payloads.
   Implementations SHOULD defend against malformed archives, zip bombs,
   polyglot files, decompression attacks, path traversal in embedded
   filenames, and parser-specific exploit inputs.  BlobBag processing
   SHOULD apply bounded resource policies for maximum object size,
   expansion ratios, page counts, recursion depth, and extraction time.

5.4.  Metadata and Structured Content Injection

   Titles, filenames, annotations, extracted fields, and table content
   MAY contain attacker-controlled strings.  Implementations SHOULD
   treat these values as untrusted input when rendering search results,
   constructing queries, invoking downstream tools, or generating
   prompts for language models.

5.5.  Cross-Tenant Trust Boundaries

   Shared indexing, storage, or enrichment infrastructure can create
   cross-tenant leakage risks if document-processing metadata is reused
   outside its intended scope.  Implementations SHOULD isolate tenant
   data paths, avoid sharing authorization context across tenants, and
   ensure that cached semantic or structured outputs are keyed by both
   doc-id and the relevant ownership boundary.

Rickert                  Expires 24 January 2027                [Page 9]
Internet-Draft             PipeStream DocProc                  July 2026

6.  IANA Considerations

   This document requests no IANA actions.

6.1.  Profile Identification

   This profile is identified by the case-sensitive string DOCPROC.
   Implementations MAY advertise this identifier in out-of-band
   configuration, capability metadata, or application-specific routing
   tables.

   The initial profile schema version defined by this document is 1.
   Receivers that do not support the advertised profile-version in a
   PipeDoc payload SHOULD reject that payload at the application layer
   rather than attempting best-effort interpretation.

7.  Normative References

   [PIPESTREAM]
              Rickert, K., "PipeStream: A Recursive Entity Streaming
              Protocol for Distributed Processing over QUIC", Work in
              Progress, Internet-Draft, draft-krickert-pipestream-03,
              July 2026, <https://datatracker.ietf.org/doc/draft-
              krickert-pipestream/>.

   [RFC2119]  Bradner, S., "Key words for use in RFCs to Indicate
              Requirement Levels", BCP 14, RFC 2119,
              DOI 10.17487/RFC2119, March 1997,
              <https://www.rfc-editor.org/rfc/rfc2119>.

   [RFC8174]  Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC
              2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174,
              May 2017, <https://www.rfc-editor.org/rfc/rfc8174>.

   [RFC8610]  Birkholz, H., Vigano, C., and C. Bormann, "Concise Data
              Definition Language (CDDL): A Notational Convention to
              Express Concise Binary Object Representation (CBOR) and
              JSON Data Structures", RFC 8610, DOI 10.17487/RFC8610,
              June 2019, <https://www.rfc-editor.org/rfc/rfc8610>.

Appendix A.  Complete CDDL Schema

   This appendix provides the profile's consolidated CDDL [RFC8610]
   definitions.

   The file-storage-reference and encryption-metadata types are defined
   by reference in PipeStream Core Appendix C and are reused here
   without modification.

Rickert                  Expires 24 January 2027               [Page 10]
Internet-Draft             PipeStream DocProc                  July 2026

   pipe-doc = {
     profile-version: uint,
     doc-id: tstr,
     entity-id: uint,
     ? search-metadata: search-metadata,
     ? blob-bag: blob-bag,
     ? semantic-result: semantic-processing-result,
     ? layer2-payload: layer2-payload,
     ? custom-entity: any,
     ? ownership: ownership-context,
   }

   search-metadata = {
     ? title: tstr,
     ? keywords: [* tstr],
     ? description: tstr,
     ? custom-fields: { * tstr => tstr },
   }

   blob-bag = blob / blobs

   blobs = {
     blobs: [* blob],
   }

   blob = {
     blob-id: tstr,
     ? drive-id: tstr,
     content: bstr / file-storage-reference,
     ? mime-type: tstr,
     ? filename: tstr,
     ? size-bytes: int,
     ? checksum: tstr,
     ? checksum-type: checksum-type,
   }

   checksum-type = &(
     unspecified: 0,
     md5: 1,
     sha1: 2,
     sha256: 3,
     sha512: 4,
   )

   semantic-processing-result = {
     ? chunks: [* semantic-chunk],
     ? chunking-strategy: tstr,
     ? processing-metadata: { * tstr => tstr },

Rickert                  Expires 24 January 2027               [Page 11]
Internet-Draft             PipeStream DocProc                  July 2026

   }

   semantic-chunk = {
     chunk-id: tstr,
     ? chunk-number: int,
     ? embedding-info: chunk-embedding,
     ? metadata: { * tstr => any },
     ? annotations: [* nlp-annotation],
   }

   chunk-embedding = {
     text-content: tstr,
     ? vector: [* float],
     ? model-id: tstr,
     ? original-char-start-offset: int,
     ? original-char-end-offset: int,
   }

   nlp-annotation = {
     type: tstr,
     label: tstr,
     ? start-offset: int,
     ? end-offset: int,
     ? confidence: float,
     ? attributes: { * tstr => tstr },
   }

   parsed-metadata = {
     parser-id: tstr,
     ? fields: { * tstr => any },
     ? tables: [* table-data],
     ? raw-output: tstr,
   }

   layer2-payload = {
     ? parsed-metadata: { * tstr => parsed-metadata },
     ? structured-data: any,
   }

   table-data = {
     table-id: tstr,
     ? headers: [* tstr],
     ? rows: [* table-row],
   }

   table-row = {
     cells: [* tstr],
   }

Rickert                  Expires 24 January 2027               [Page 12]
Internet-Draft             PipeStream DocProc                  July 2026

   ownership-context = {
     ? tenant-id: tstr,
     ? owner-id: tstr,
     ? acl: [* tstr],
   }

Appendix B.  Example Processing Patterns

   This appendix is non-normative.

B.1.  Text Extraction

   A text-extraction stage transforms binary document content into
   textual or layout-aware intermediate representations.  Typical
   outputs include page text, OCR results, or format-specific structural
   markup.

B.2.  NLP Enrichment

   An enrichment stage adds semantic metadata such as chunking,
   embeddings, named entities, classifications, or relation annotations.
   These results are commonly encoded in SemanticLayer payloads.

B.3.  Structured Table Extraction

   A table-extraction stage identifies structured tabular regions and
   emits normalized table representations suitable for indexing or
   analytics.

B.4.  Image Processing

   An image-processing stage derives metadata or features from document
   images, such as OCR overlays, captions, detections, or classification
   results.

B.5.  Example Sink Patterns

   Common sink patterns include search indexing, archival persistence,
   and workflow notification.  Backend-specific integrations such as
   particular search engines, object stores, or message buses are
   deployment choices, not protocol requirements of this profile.

Author's Address

   Kristian Rickert
   PipeStream AI
   Email: kristian.rickert@pipestream.ai

Rickert                  Expires 24 January 2027               [Page 13]