Network Working Group S. Dikshit Internet-Draft Aruba Networks, HPE Intended status: Informational M. Srivastava Expires: 28 January 2027 Hewlett Packard Enterprise 28 July 2026 Network Management and OAM Considerations for IP in Deep Space draft-dikshit-tiptop-oam-considerations-01 Abstract [I-D.ietf-tiptop-usecase] and [I-D.ietf-tiptop-ip-architecture] describe key characteristics, use cases, requirements, and an IP architecture for deep-space (lunar, Mars, and beyond) surface and orbital-relay networking, characterized by long, variable, asymmetric propagation delay (single-digit to tens of minutes one-way), scheduled/intermittent connectivity windows, and severely constrained link capacity. Section 8.2 of [I-D.ietf-tiptop-ip-architecture] addresses the configuration- management plane for this environment (NETCONF, RESTCONF, and SNMP transport selection and RTT-adjusted client timeouts), but neither document addresses the fault and performance dimensions of Operations, Administration, and Maintenance (OAM): fault detection, performance measurement, and reachability verification. This document identifies why conventional terrestrial OAM techniques (active round-trip probing such as ICMP Echo, BFD, or TWAMP; assumption of continuous connectivity) do not transfer directly to this environment, and proposes candidate adapted approaches -- passive, store-and-forward telemetry batched to contact windows; on-board local health self-diagnosis with delayed reporting; and confidence-interval-based reachability assessment in place of binary up/down status -- as a starting point for OAM discussion within the TIPTOP working group. This document explicitly does not propose use of the Bundle Protocol or DTN architecture, both of which are out of scope for TIPTOP per its charter; where CCSDS space-data-system telemetry conventions are mentioned, it is as informative background only. Status of This Memo This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79. Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/. Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress." This Internet-Draft will expire on 28 January 2027. Copyright Notice Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved. This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/ license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Dikshit & Srivastava Expires 28 January 2027 [Page 1] Internet-Draft TIPTOP OAM Considerations July 2026 Table of Contents 1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . 2 1.1. Requirements Language . . . . . . . . . . . . . . . . . . 3 2. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . 3 3. Why Terrestrial OAM Does Not Transfer Directly . . . . . . . 3 3.1. Active Round-Trip Probing . . . . . . . . . . . . . . . 3 3.2. Continuous-Connectivity Assumption . . . . . . . . . . . 4 3.3. Binary Reachability Semantics . . . . . . . . . . . . . 4 4. Candidate Adapted OAM Approaches . . . . . . . . . . . . . . 5 4.1. Passive, Store-and-Forward Telemetry . . . . . . . . . . 5 4.2. On-Board Local Health Self-Diagnosis . . . . . . . . . . 5 4.3. Confidence-Interval-Based Reachability . . . . . . . . . 6 4.4. Contact-Window-Aware Fault Correlation on the Ground Segment . . . . . . . . . . . . . . . . . . . . . . . . . 6 4.5. Complementarity with Architecture Section 8.2 . . . . . . 7 5. Explicitly Out of Scope . . . . . . . . . . . . . . . . . . 7 6. Relationship to Existing Work . . . . . . . . . . . . . . . . 7 7. Security Considerations . . . . . . . . . . . . . . . . . . . 8 8. IANA Considerations . . . . . . . . . . . . . . . . . . . . . 8 9. Acknowledgements . . . . . . . . . . . . . . . . . . . . . . 8 10. References . . . . . . . . . . . . . . . . . . . . . . . . . 8 10.1. Normative References . . . . . . . . . . . . . . . . . . 8 10.2. Informative References . . . . . . . . . . . . . . . . . 9 Appendix A. Changes from -00 . . . . . . . . . . . . . . . . . . 9 Authors' Addresses . . . . . . . . . . . . . . . . . . . . . . . 9 1. Introduction [I-D.ietf-tiptop-usecase] catalogs key characteristics, use cases, and requirements for IP networking in deep space: propagation delay ranging from single-digit minutes (Earth-Moon) to tens of minutes one-way (Earth-Mars, depending on orbital geometry), scheduled and intermittent connectivity governed by orbital mechanics and relay visibility windows, severe and asymmetric bandwidth constraints, and the need to interoperate with existing space-agency ground segment infrastructure. [I-D.ietf-tiptop-ip-architecture] builds an IP architecture addressing these characteristics. Section 8.2 of [I-D.ietf-tiptop-ip-architecture] addresses the management plane to the extent of protocol and transport selection: NETCONF [RFC6241] and RESTCONF [RFC8040] carried over QUIC or HTTP/3 with appropriate transport parameters, NETCONF client timeouts adjusted to the expected round-trip time, and SNMP [RFC1157] identified as usable unmodified given suitable timeout configuration. It further observes that a configuration change may take hours to reach a spacecraft, and that a requested value may already be stale by the time it reaches the requestor. That treatment is, however, confined to configuration management. It does not extend to the fault and performance dimensions of OAM: neither document discusses how an operator (terrestrial or, in the future, non-terrestrial) is meant to detect faults, measure performance, or verify reachability in this environment. This gap matters: every terrestrial IP OAM tool in common use (ICMP Echo Request/Reply, BFD [RFC5880], TWAMP [RFC5357], even basic traceroute) is designed around an implicit assumption that a round trip completes in well under a second and that the path is continuously available to probe. Applying such tools unmodified over a link with minutes of one-way delay and scheduled connectivity windows would, at best, produce results so stale as to be operationally useless (a "down" indication based on a probe sent twenty minutes ago), and at worst consume a meaningful fraction of an already severely constrained link's capacity and contact-window duration on probe traffic rather than payload. This document is intended as a discussion starter for the TIPTOP working group, not a complete specification. It identifies the specific ways terrestrial OAM assumptions break down (Section 3) and sketches candidate adapted approaches (Section 4) that the working group may wish to develop into a full OAM requirements document, analogous to how [I-D.ietf-cats-oam-fw] was scoped out as a distinct framework once the base CATS architecture existed. 1.1. Requirements Language The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here. 2. Terminology This document uses terminology from [I-D.ietf-tiptop-usecase] and [I-D.ietf-tiptop-ip-architecture], including "contact window" (a scheduled period during which a link between two deep-space nodes, or between a deep-space node and a ground station, is expected to be available) and "relay node" (an orbital or surface node forwarding traffic between a deep-space node and Earth). 3. Why Terrestrial OAM Does Not Transfer Directly 3.1. Active Round-Trip Probing ICMP Echo, BFD, and TWAMP all measure liveness/performance by sending a probe and waiting for a corresponding reply, on a timescale of milliseconds to a few seconds. Over an Earth-Mars path, one round trip alone can take 8 to 48 minutes depending on orbital geometry, before accounting for any queuing or relay processing delay. A probe-and-wait model at this timescale is still theoretically possible but operationally marginal: any fault-detection or performance-measurement system built on it would have a detection latency measured in tens of minutes at best, and would consume contact-window time and link capacity that is already the scarcest resource in the environment described by [I-D.ietf-tiptop-usecase]. 3.2. Continuous-Connectivity Assumption Terrestrial OAM tools generally assume a monitored path is available to probe at any time; an unreachable probe target is itself the fault signal. In deep space, links are scheduled and intermittent by design (orbital visibility, power/thermal constraints, competing mission priorities for relay time) -- "currently unreachable because outside the contact window" and "currently unreachable because faulty" are both common and must be distinguished. An OAM approach that cannot tell these apart will generate false fault indications on every scheduled gap. 3.3. Binary Reachability Semantics Conventional OAM reports reachability as a binary up/down state as of the last successful/failed probe. Given the delay and intermittency described above, a reachability report is unavoidably stale by the time it is observed on the other end of the link; a binary status does not communicate how stale, nor how confident the observer should be that the reported state still holds. An approach that reports a probabilistic or confidence-interval-based status (Section 4.3), rather than a single stale boolean, better matches the actual information available. 4. Candidate Adapted OAM Approaches 4.1. Passive, Store-and-Forward Telemetry Rather than active probe/reply exchanges, nodes SHOULD locally accumulate OAM-relevant telemetry (interface counters, queue depths, local error events) continuously, and transmit accumulated telemetry as a batch at the next scheduled contact window, rather than attempting real-time exchange. This trades detection latency (bounded by contact-window interval rather than round-trip time, which is a strictly better bound whenever the contact-window interval is shorter than a round trip, and a known, predictable bound in all cases since contact windows are scheduled) for dramatically reduced probe traffic overhead. This approach draws conceptually on store-and-forward telemetry conventions long used in CCSDS space data systems [CCSDS-133], referenced here purely as informative prior art, not as a proposal to adopt CCSDS protocols or the Bundle Protocol/DTN architecture (see Section 5). 4.2. On-Board Local Health Self-Diagnosis Because round-trip active probing is impractical (Section 3.1), fault detection should shift toward local self-diagnosis: a node monitors its own interfaces, queues, and processing health continuously and independently, raising a local alarm the moment a local threshold is crossed, and includes that alarm (with its local timestamp) in the next batch of telemetry sent at the next contact window (Section 4.1). This is the same principle behind on-board health monitoring in spacecraft subsystems generally, applied here specifically to the IP networking layer. 4.3. Confidence-Interval-Based Reachability Rather than reporting a binary up/down state, an OAM status report SHOULD include an explicit "as-of" timestamp and, where derivable, an expected next-observation time based on the next scheduled contact window, so that a consumer of the report can compute its own confidence that the reported state still holds, rather than treating a stale report as current. A gap in expected telemetry that extends beyond the next scheduled contact window (rather than a single missed probe) is a more meaningful fault signal in this environment than any single probe timeout. 4.4. Contact-Window-Aware Fault Correlation on the Ground Segment Fault correlation and root-cause analysis are expected to happen primarily on the Earth (ground segment) side, where processing power and continuous connectivity to other ground systems are not constrained the way they are on deep-space nodes. The ground segment SHOULD correlate telemetry batches (Section 4.1) and local alarms (Section 4.2) from multiple nodes/relays against the known contact-window schedule to distinguish "this gap in reporting is scheduled and expected" from "this gap in reporting exceeds the scheduled window and is a fault," rather than expecting each deep-space node to perform this correlation itself. 4.5. Complementarity with Architecture Section 8.2 Section 8.2 of [I-D.ietf-tiptop-ip-architecture] establishes the configuration-management plane for deep-space IP nodes. The approaches in this section are intended to complement, not replace, that treatment. Three specific points of contact are worth calling out, each of which the authors offer as candidate additions to a future revision of the architecture document. Shared contact-window scheduling. Configuration operations carried over NETCONF or RESTCONF and telemetry batches (Section 4.1) both contend for the same scarce contact windows and the same constrained link capacity. Treating them as independent traffic classes risks a configuration push and a telemetry drain colliding within a single window. A shared scheduling abstraction -- in which management-plane and OAM-plane traffic are jointly planned against the known contact schedule, with explicit relative priority -- is preferable to per-protocol timeout tuning alone. Explicit staleness annotation rather than timeout extension alone. Section 8.2 observes that a value returned in response to a request "may not be current or expired," and recommends adjusting NETCONF client timeouts to the expected round-trip time. Adjusting timeouts prevents the client from erroneously declaring failure, but it does not tell the operator how stale the delivered value actually is. This is the same problem that Section 4.3 addresses for reachability. The authors suggest that operational data retrieved from a deep-space node SHOULD carry an explicit observation timestamp and, where the node can supply one, an indication of the interval over which the value is expected to remain valid, so that the ground segment can reason about staleness directly rather than inferring it from request latency. Contact-window-aware notification buffering. Section 8.2 does not discuss YANG notification subscriptions [RFC8639] [RFC8641]. An on-change subscription cannot deliver during a scheduled connectivity blackout, and an on-change event generated mid- blackout will either be lost or delivered arbitrarily late with no indication that this occurred. A subscription mechanism usable in this environment requires on-board buffering of notifications across the blackout, replay at the next contact window, and an explicit indication to the receiver that the delivered events are historical rather than current. This is the notification-plane analogue of the batched telemetry described in Section 4.1, and the authors believe it is a concrete gap in the current architecture text. 5. Explicitly Out of Scope Per the TIPTOP working group charter, the Bundle Protocol and Delay/Disruption-Tolerant Networking (DTN) architecture [RFC4838] are out of scope, as are LEO/MEO/GEO satellite-constellation IP networking topics. This document does not propose adopting BP/DTN mechanisms; references to CCSDS conventions (Section 4.1) are included solely as informative background on store-and-forward telemetry design in comparable environments, consistent with the working group's existing practice of treating space-agency ground- segment conventions as background rather than normative dependency. 6. Relationship to Existing Work * [I-D.ietf-tiptop-usecase] and [I-D.ietf-tiptop-ip-architecture] are the two existing WG documents this document proposes extending with an OAM discussion. * Section 8.2 of [I-D.ietf-tiptop-ip-architecture] specifies the configuration-management plane (NETCONF, RESTCONF, and SNMP transport selection and timeout tuning). This document does not revisit or contradict that treatment; it addresses the fault and performance dimensions of OAM, which Section 8.2 does not cover. Section 4.5 identifies three specific points at which the two interact. * [I-D.ietf-cats-oam-fw] is cited as a structural precedent for how another working group scoped an OAM framework as a distinct document once its base architecture existed; no technical content is shared, since CATS OAM assumes continuous, low- latency connectivity that does not hold here. * [RFC5880] (BFD) and [RFC5357] (TWAMP) are cited in Section 3.1 as examples of terrestrial active-probing OAM tools whose underlying round-trip assumption motivates this document. 7. Security Considerations This document proposes no new protocol mechanisms and introduces no new security considerations beyond those already applicable to [I-D.ietf-tiptop-ip-architecture]. Batched telemetry (Section 4.1) sent at contact windows SHOULD be integrity-protected and, where confidentiality of operational telemetry matters (e.g., commercial relay operators), encrypted, consistent with the general security posture already expected of the deep-space IP architecture. Local health self-diagnosis (Section 4.2) alarms, once relayed to the ground segment, become part of that same telemetry stream and inherit the same protection. 8. IANA Considerations This document has no IANA actions. 9. Acknowledgements Thanks to the authors of [I-D.ietf-tiptop-usecase] and [I-D.ietf-tiptop-ip-architecture] for the foundational documents this discussion builds on. 10. References 10.1. Normative References [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, March 1997. [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, May 2017. [I-D.ietf-tiptop-usecase] "IP in Deep Space: Key Characteristics, Use Cases and Requirements", Work in Progress, draft-ietf-tiptop- usecase-02, 6 July 2026. [I-D.ietf-tiptop-ip-architecture] "An Architecture for IP in Deep Space", Work in Progress, draft-ietf-tiptop-ip-architecture-01, 5 July 2026. 10.2. Informative References [RFC1157] Case, J., Fedor, M., Schoffstall, M., and J. Davin, "A Simple Network Management Protocol (SNMP)", RFC 1157, May 1990. [RFC4838] Cerf, V., Burleigh, S., Hooke, A., Torgerson, L., Durst, R., Scott, K., Fall, K., and H. Weiss, "Delay-Tolerant Networking Architecture", RFC 4838, April 2007. [RFC5880] Katz, D. and D. Ward, "Bidirectional Forwarding Detection (BFD)", RFC 5880, June 2010. [RFC5357] Hedayat, K., Krzanowski, R., Morton, A., Yum, K., and J. Babiarz, "A Two-Way Active Measurement Protocol (TWAMP)", RFC 5357, October 2008. [RFC6241] Enns, R., Bjorklund, M., Schoenwaelder, J., and A. Bierman, "Network Configuration Protocol (NETCONF)", RFC 6241, June 2011. [RFC8040] Bierman, A., Bjorklund, M., and K. Watsen, "RESTCONF Protocol", RFC 8040, January 2017. [RFC8639] Voit, E., Clemm, A., Gonzalez Prieto, A., Nilsen-Nygaard, E., and A. Tripathy, "Subscription to YANG Notifications", RFC 8639, September 2019. [RFC8641] Clemm, A. and E. Voit, "Subscription to YANG Notifications for Datastore Updates", RFC 8641, September 2019. [I-D.ietf-cats-oam-fw] Fu, W., Xiong, Q., Du, Z., Liu, P., and X. Li, "Computing-Aware Traffic Steering (CATS) Operations, Administration, and Maintenance (OAM) Framework", Work in Progress, draft-ietf-cats-oam-fw-01, 21 July 2026. [CCSDS-133] Consultative Committee for Space Data Systems, "Space Packet Protocol", CCSDS 133.0-B-2, informative reference only. Appendix A. Changes from -00 * Corrected the characterization of the existing working group documents. The -00 revision stated that neither [I-D.ietf-tiptop-usecase] nor [I-D.ietf-tiptop-ip-architecture] addressed network management or OAM. That was inaccurate: Section 8.2 of [I-D.ietf-tiptop-ip-architecture] specifies the configuration-management plane. The Abstract and Section 1 now cite Section 8.2 explicitly and narrow the gap this document claims to the fault and performance dimensions of OAM. * Added Section 4.5, describing three concrete points of complementarity with Section 8.2: shared contact-window scheduling for management and OAM traffic, explicit staleness annotation on retrieved operational data rather than timeout extension alone, and contact-window-aware buffering and replay of YANG notification subscriptions. * Added a corresponding bullet to Section 6. * Added [RFC1157], [RFC6241], [RFC8040], [RFC8639], and [RFC8641] as informative references. Authors' Addresses Saumya Dikshit Aruba Networks, Hewlett Packard Enterprise Email: saumya.dikshit@hpe.com Mukul Srivastava Hewlett Packard Enterprise Email: mukul.srivastava@hpe.com