VCF Operations and Observability in VCF 9.1 A Day 0/1/2 Field Guide
VCF Operations is now the centre of gravity for the whole platform not just monitoring, but lifecycle, capacity, compliance, and diagnostics. Here’s how to design, deploy, and operate it.
VCF Operations has moved from being a monitoring layer above the platform to being the operational centre of gravity of the platform itself. Lifecycle, capacity, licensing, compliance, diagnostics, fleet management, identity they all run through VCF Operations now. For architects, this changes what the observability conversation is about: it’s no longer “which metrics do we collect,” it’s “how does VCF Operations become the single pane that runs the platform.”
This piece walks through the Day 0 design decisions, Day 1 deployment patterns, and Day 2 operational use of VCF Operations in VCF 9.1.
Day 0 Designing the Observability Layer
The Day 0 design conversation for VCF Operations starts with scope: which capabilities are in, which are out, and where the integration points sit.
• Fleet scope single VCF Operations instance can govern multiple VCF instances. Decide whether to centralise or federate based on regulatory boundaries, network isolation requirements, and operational ownership.
• Appliance topology High Availability model deploys 3× VCF Operations nodes, 3× VCF Operations for Logs nodes. Production environments should default to HA; Simple model is acceptable only for non-production.
• Management Packs to install Diagnostics MP is a core install. Custom MPs (hardware vendor, storage array, third-party) added based on environment.
• Diagnostics Content Pack vs Management Pack in VCF 9.1, Diagnostics Content Pack ships separately from the Management Pack so findings update on a faster cadence than the MP releases. Architects need to know both exist.
• Integration boundaries SIEM (Splunk, Sentinel, QRadar) via Syslog; ITSM (ServiceNow) via VCF Operations APIs; cost intelligence (Broadcom Business Services console) integrated natively.
The Real-Time Operational Observability story in VCF 9.1 is the architectural shift: 2-second metric resolution with PromQL queryability. This brings VCF Operations into parity with cloud-native observability platforms. The HLD implication: SLO/SLI definitions become viable for traditional infrastructure, not just modern apps. Architects who have only thought of SLOs as a containerised concept can extend the discipline to VMs and infrastructure services.
Day 1 Deploying VCF Operations Correctly
Day 1 deployment of VCF Operations in VCF 9.1 follows the new VCF Installer workflow, which simplifies the appliance deployment but introduces dependencies architects need to plan for.
• Resource reservations HA model needs sufficient capacity in the Management Domain for 3× VCF Operations + 3× Operations for Logs + 3× NSX Manager + 3× VCF Automation. The Planning and Preparation Workbook calculates this.
• DNS records lower-case FQDNs only. VCF Operations DNS record with even a single upper-case letter will fail the integration. This is the most common Day 1 failure mode.
• Diagnostics installation install both the Diagnostics Management Pack and Diagnostics Content Pack. The MP releases with VCF Operations updates; the CP releases independently with new findings.
• Management Pack baseline install vSphere MP, NSX MP, vSAN MP at minimum. Each surfaces domain-specific findings into VCF Operations’ unified Active Findings view.
• Log Assist enablement enable Log Assist for most VCF components so log-based findings surface alongside property-based findings.
A note on Diagnostics evolution by version: VCF Operations 8.18.x had ~250 property findings and 30 log findings, mostly build/version rules; 9.0.x added Skyline Advisor-equivalent rules totalling ~600 property findings + 100 log findings + hybrid findings; 9.1.x splits the MP and CP for faster content updates. Architects supporting customers on vSphere 8.x / NSX 4.x can install VCF Operations 9.1 and get the latest Diagnostics findings without forcing a stack upgrade.
The Day 1 dashboarding decision: do you build custom dashboards or use the shipped views? Out-of-the-box dashboards cover the common operational views (capacity, performance, compliance). Custom dashboards are appropriate when you have specific tenant views, regulatory reporting requirements, or business KPIs to surface. Keep custom dashboards minimal at Day 1; let operations evolve the view based on actual usage patterns.
Day 2 Operating with VCF Operations as the Centre of Gravity
Day 2 operations through VCF Operations cover capabilities that previously required multiple tools. The architectural value is single-pane control across the lifecycle.
Real-Time Metrics and Active Findings
2-second metric resolution with PromQL queryability changes troubleshooting. Performance investigations that used to require packet captures or external monitoring can now query telemetry directly from VCF Operations. Active Findings prioritise the signal: instead of dashboards full of yellow alerts, the platform surfaces a curated set of findings that need action, with remediation guidance per finding.
Capacity and Cost
Effective Capacity View in vSAN integrates with VCF Operations Capacity. The Operations capacity insights also surface workload candidates for NVMe Memory Tiering automatically the architect doesn’t have to hunt for them. Cost intelligence connects to the Broadcom Business Services console for licence consumption and to internal billing integration via APIs.
Compliance and Continuous Posture
Continuous Compliance Enforcement via Advanced Cyber Compliance (ACC) 9.1 surfaces drift from PCI DSS and VCF security baselines as Active Findings. The compliance model shifts from periodic audit to continuous posture. Audit evidence becomes a byproduct of the platform operating normally, not a special event.
Lifecycle and Fleet Governance
Fleet Update Service, Quick Patch orchestration, ESX Live Patch scheduling, backup configuration all driven from VCF Operations. Workload domain creation, cluster expansion, host commissioning are now in vCenter; VCF Operations governs the fleet-wide policy and orchestration layer above.
The customer testimony in the May 2026 Broadcom operations post is instructive: Japan Racing Association uses VCF Operations to gain full visibility into usage trends and performance forecasts, identifying issues before they impact mission-critical race day operations. A Broadcom customer survey (n=44) showed an average 51% reduction in infrastructure management time using VCF 9. The architectural story holds up in the field.
Architect’s Takeaway
VCF Operations is no longer a monitoring product bolted to the platform it is the platform’s operational control plane. The Day 0 design decisions (fleet scope, HA topology, Management Packs) and Day 1 deployment hygiene (DNS, integration points, Diagnostics CP) determine how effective Day 2 operations can be. For VCF architects, the operational model maturity question is now answerable through one tool, which means staffing models and operational runbooks should be designed around it. Resist the temptation to layer external observability tools on top; the closer you can get to VCF Operations as the single pane, the simpler the operational story becomes.
Sources
• Broadcom Scale, Simplify, and Secure Your Private Cloud Operations with VCF 9.1
• Broadcom Diagnostics for VMware Cloud Foundation (VCF) 9.1 with Old Versions of VCF Components
• Broadcom 10 VCF 9 Enhancements: Simplifying Your Day 2 Operations
• Broadcom Announcing VCF 9.1: Modern Private Cloud Built for Efficiency and Resilience
• Broadcom VCF 9.1: The Secure, Cost-Effective Private Cloud Platform for Production AI