Skip to content
INTERNET INFRASTRUCTURE & NETWORKS

Redesigning Infrastructure for the AI Era: Key Takeaways from APNIC 62’s Technical Session on Network AI and Automation

As modern networks rapidly evolve to support the soaring computational demands of artificial intelligence workloads, network operators are finding themselves forced to rethink both their physical infrastructure and the software-driven methods used to manage increasingly complex environments. This paradigm shift took center stage at APNIC 62 during Technical Session 2, aptly themed "AI in Networks." Industry experts gathered to explore how recent advances in data center architecture, machine learning models, agentic AI frameworks, and operational automation are fundamentally reshaping how networks are designed, scaled, and day-to-day managed.

The presentations delivered throughout the session spanned a wide array of critical topics, ranging from next-generation optical fabrics and autonomous network management workflows to predictive fault detection and AI-assisted geofeed authoring. Together, these discussions highlighted both the immense operational opportunities and the complex technical challenges inherent in building and operating robust network infrastructure at scale in the modern era.

From Fat Tree to Jellyfish: Rewiring India’s Networks for the AI Era

Opening the session, Lalit Singh Chowdhary, Chief Technology and Innovation Officer at Lightstorm Telecom Connectivity Pvt Ltd, examined the often-overlooked physical and optical infrastructure that underpins modern data center design. Providing a striking perspective on hardware scaling, Chowdhary noted that back in 2019, approximately 200 million Ethernet ports were sold globally for use inside switching fabrics. Today, that figure has skyrocketed to nearly one billion ports, with an astonishing 40 percent operating at speeds of 400Gbps or faster. In fact, there are now more active high-speed network ports being deployed than there are human babies being born.

This explosive growth forces a critical evaluation of traditional data center topologies. The concept of the "fat tree" architecture originated back in 1985 through the pioneering research of Charles Leiserson at the Massachusetts Institute of Technology. In a traditional fat tree layout, as traffic moves upward through the switching hierarchy, available bandwidth increases proportionally to prevent network bottlenecks. This foundational design underpins the classic switch-edge-aggregation-core architecture, commonly referred to as the spine-and-leaf model. Within the core of such a network, oversubscription is completely unacceptable, prompting operators to design these layers with a strict 1:1 capacity ratio, while tolerating some degree of oversubscription at the network edge.

Ultimately, Chowdhary explained, traditional data center architecture relies heavily on three core factors: the number of equal-cost paths available, the time required for data to traverse those paths, and the level of oversubscription an operator can safely accept while maintaining high reliability standards, such as five nines availability. Historically, these networks were meticulously designed to handle front-end web workloads, reflecting the capacity and metro-area latency requirements of past decades.

To address these limitations, Lightstorm utilizes a three-spine, fully meshed interconnect architecture, where network growth occurs primarily at the leaf layer. Under this configuration, any connected data center can reach another within a maximum of two hops. Interconnects are configured as dedicated wavelengths, and live traffic is distributed efficiently across the three-spine design to achieve five nines availability—equivalent to roughly 26 seconds of downtime per month. If a specific path fails, traffic can be rapidly redistributed without requiring large-scale, disruptive workload migrations between disparate data centers.

However, operating at this scale introduces persistent economic and technical hurdles. Network operators are under immense pressure to transition their legacy 10Gbps and 40Gbps line cards to much faster 100Gbps and 400Gbps platforms, a transition that can prove financially and operationally difficult for customers accustomed to older rack and interconnect cost models. Furthermore, growing application complexity, data sharding requirements, and modern fan-out traffic patterns make intelligent workload placement increasingly challenging.

To overcome these structural roadblocks, Chowdhary introduced an alternative design philosophy known as the "jellyfish" architecture. Developed collaboratively by researchers at the University of Illinois Urbana-Champaign and HP Labs, the jellyfish model completely replaces rigid, structured tiers with a conceptually random network topology that features no discernible core. This innovative design relies on colourless, directionless, and contentionless reconfigurable optical add-drop multiplexer, or CDC-ROADM, technology, significantly reducing path costs through advanced wavelength switching. The underlying wavelength-switching plane remains fully transparent, allowing any wavelength to travel in virtually any direction without encountering hardware contention. If a required wavelength is already occupied by existing traffic, the system can dynamically recolor and reroute the traffic on the fly.

This fundamental architectural shift is ultimately driven by industry projections anticipating a tenfold increase in overall capacity demand. Future network growth will likely be constrained primarily by local power availability, requiring computational workloads to migrate dynamically to geographic locations where power is abundant, while network capacity follows suit. By utilizing a fully integrated optical interconnect operating under a unified management framework, operators can maintain rigorous reliability standards even as the underlying physical topology transforms.

AI in Network Operations and Agentic Workflows

Transitioning from physical topologies to software-driven management, Akhilesh Thakur of HPE Networking outlined a compelling vision for fully autonomous networks requiring minimal manual intervention, which would drastically reduce the operational burden traditionally associated with reactive troubleshooting.

Previous corporate initiatives aimed at introducing artificial intelligence into network management frequently stumbled due to a fundamental lack of sufficient, high-quality data. Today, however, network operators have access to vastly richer telemetry datasets spanning active network elements, control and data planes, underlying infrastructure platforms, critical business systems, and external threat intelligence sources.

Thakur placed significant emphasis on the concept of agentic AI, a sophisticated approach that breaks down complex operational problems into smaller, manageable actions while carefully maintaining contextual awareness through a retrieval-augmented generation knowledge base. To facilitate communication between these AI agents and operational tools, the Model Context Protocol serves as a standardized intermediary interface. Rather than requiring developers to build and maintain direct integrations with disparate APIs, databases, or NetConf implementations, MCP acts as a consistent abstraction layer. This architectural choice provides agentic systems with a uniform operational view across multi-vendor platforms while significantly cutting down integration complexity.

During his presentation, Thakur demonstrated a live implementation of Junos MCP. Following HPE’s high-profile acquisition of Juniper Networks, robust Junos routing and switching platforms can now be seamlessly integrated into large-scale AI and data center environments. Advanced large language models like Claude can interact natively with these network environments through Junos MCP, leveraging feature introspection to identify available network capabilities dynamically and translate high-level natural language requests into precise, executable operational workflows.

Demonstrated use cases included generating comprehensive network audits through conversational interfaces, as well as integrating MCP with external threat intelligence sources such as Cloudflare Radar and RIPEstat to rapidly detect and investigate complex BGP route hijacking incidents. Thakur also showcased higher-order MCP models capable of synthesizing telemetry data from hosts, routers, and switches simultaneously to pinpoint root causes during performance degradation events.

Despite these powerful capabilities, Thakur cautioned that fully autonomous closed-loop operation remains a long-term objective, and current deployments sensibly maintain a strict human-in-the-loop oversight model. Automated generation of Methods of Procedure, for instance, benefits greatly from combining LLM inference with retrieval-augmented generation, producing upgrade procedures that align closely with intended operational outcomes while minimizing unnecessary steps and mitigating risks such as hallucinations, credential exposure, and model drift.

AI-driven networking and infrastructure design at APNIC 62 | APNIC Blog

Live Machine Learning Solutions for Broadband Network AIOps

Delving into the practical application of machine learning for telecommunications, Dr. Girish Saraph of Vegayan Systems Pvt. Ltd. addressed the widening gap between rapid network growth and the limited engineering resources available to small and medium-sized enterprises and traditional Network Operations Centres. With consumer expectations for broadband responsiveness higher than ever, network operators must prioritize critical issues effectively and identify potential hardware failures well before they impact end users.

To illustrate the severity of the operational scaling problem, Dr. Saraph presented a case study of a large broadband network comprising 10 million active customer endpoints distributed across optical network terminal and gigabit passive optical network infrastructure. This massive environment routinely generated 10 million distinct alarms daily, translating to approximately 200 active, high-priority issues every single day. Historically, the network required an average of four to six hours from initial diagnosis to final remediation—a workload volume that far exceeded the capacity of the available NOC staff.

Because missing a major infrastructure outage can trigger severe service-level agreement financial penalties, automated prioritization is an absolute necessity. Vegayan Systems developed an MLOps platform designed to act as an automated network expert, ingesting historical fault and outage records spanning several years to train multiple specialist machine learning models. Each model focuses on a narrow problem domain, processing live telemetry to generate predictive insights and intelligent prioritization recommendations.

When deployed in ONT and PON environments, the platform successfully identified subscriber access ports with a history of recurring, intermittent faults. In the network studied, roughly 7.5 percent of customer access ports fell into this high-risk category, and the machine learning system achieved an impressive 74 percent hit rate in predicting problematic ports ahead of time. By identifying these degradation patterns early, operators could proactively remediate physical faults before customers experienced noticeable service disruptions, thereby lowering subscriber churn.

Applying the exact same predictive methodology to GPON aggregation nodes yielded similarly strong operational results. Approximately 5.8 percent of aggregation nodes in the test network were classified as high risk, with the predictive model demonstrating a 65 percent accuracy rate. This provided NOC engineers with a critical three- to four-day prediction window, enabling them to easily distinguish between external fiber cuts and internal hardware-related anomalies such as component overheating or voltage instability. Ultimately, this proactive approach improved overall service performance by 20 percent.

Given the staggering volume of raw alarms generated by modern broadband networks, the platform consolidates related alarms into coherent events, assigns appropriate severity rankings, and sorts incidents strictly by customer impact. This intelligent triage allows overburdened NOC teams to concentrate their energy exclusively on the most critical operational issues first, validating Dr. Saraph’s core message that operators should rigorously test multiple machine learning algorithms to select the optimal model for their unique network environment.

Authoring IP Geofeeds Using AI Tools and MCP

Shifting the focus to internet routing data hygiene, Sid Mathur of Fastah Inc. delivered an insightful presentation covering geofeed publication best practices and how modern MCP-enabled AI tools can drastically improve overall geofeed data quality. Mathur introduced the topic by sharing a common real-world frustration: an internet user whose IP address is incorrectly mapped to an entirely different economy due to inaccurate third-party geolocation databases. In most cases, frustrated content providers mistakenly direct these geolocation complaints back to the customer’s internet service provider for resolution.

To combat this widespread issue, ISPs publish structured geolocation data for their allocated IP address blocks using the standardized format defined in RFC 8805. Resembling a simple comma-separated values file, this format emerged from a clear industry and regulatory need to reliably identify the geographic location of IP address allocations. Under this model, network operators associate a specific IP prefix with an ISO 3166 economy code, a regional identifier, and a city name. Because all specific location fields are entirely optional, operators retain the flexibility to publish data at the broad economy level, the regional level, or down to the specific city level.

Furthermore, operators can explicitly publish a value of "NONE" to indicate to downstream consumers that specific address blocks should not be geolocated at all. This capability is particularly vital for core backbone infrastructure addresses and anycast network deployments, where assigning a rigid physical location to a globally distributed service can introduce routing inefficiencies and misleading mapping data.

Additional standards such as RFC 9632 establish formal discovery mechanisms within Regional Internet Registry frameworks, allowing geofeed consumers to easily locate and interpret published information. Meanwhile, newer proposals like RFC 9877 aim to represent this identical geographic data using a modern, JSON-based Registration Data Access Protocol model.

However, Mathur stressed that geofeed management must be treated as an ongoing operational process rather than a static, one-time administrative task. While smaller ISPs may update their location records infrequently, larger, dynamic networks require continuous updates as IP allocations shift. Publishing geofeeds through an RIR-linked URL places direct responsibility for data accuracy squarely on the address holder. While platforms like GitHub provide reliable hosting for these files, external database providers have no direct control over the underlying content. Because accurate geolocation inherently carries privacy implications, city-level accuracy is generally sufficient and should always be viewed as an approximate geographic indicator rather than precise coordinates.

To streamline this administrative burden, Mathur explored the integration of large language models within standard geofeed management workflows. When provided with proper RFC guidance and internal IP address management documentation, an LLM can easily generate fully compliant geofeed files and conduct automated data audits.

Demonstrating his newly developed AI skill focused specifically on geofeed management, Mathur showed how automated tools can quickly spot common errors in publicly accessible geofeeds, such as missing regional identifiers for ambiguously named cities like Frankfurt, excessive use of do-not-geolocate entries, or invalid numeric values mistakenly substituted for official ISO 3166 country codes. Operating much like a traditional software linter, these automated validation checks significantly reduce human error. Additionally, specialized Model Context Protocol servers can supply artificial intelligence agents with geographic reference data to accurately distinguish between cities sharing identical names worldwide.

Concluding his remarks, Mathur strongly recommended the GitHub-hosted CSV publication model and demonstrated how his semantic validation tooling can be integrated smoothly into everyday network operations.

Across all the diverse presentations delivered during Technical Session 2, a unifying theme emerged clearly: the growing, undeniable industry necessity for greater intelligence, operational flexibility, and deep automation throughout every layer of the network stack. Whether through revolutionary data center topologies designed specifically for massive AI workloads, autonomous software agents assisting network engineers, predictive machine learning systems forecasting hardware faults before they occur, or semantic tools ensuring the integrity of routing and geolocation data, the ultimate goal remains consistent. Modern network operators are increasingly empowered to manage staggering technological complexity with heightened confidence, resilience, and operational efficiency.

Leave a Reply

Your email address will not be published. Required fields are marked *