The VCF 9.1 Design Gaps Checklist What Architects Need to Decide Before They Build
Most VCF 9.1 design failures I've seen weren't caused by missing documentation. They were caused by a decision nobody actually made. Eighteen of them, by domain, for your next HLD review.
Broadcom's documentation for VCF 9.1 is good. The design guide, the 218-slide Designs and Solutions deck, the release notes, the whitepapers all of it holds up. But none of it can make the decisions that are specific to your estate. Every design failure I've run into on this platform generation comes back to the same thing: not a missing feature, not a documentation gap, but a decision that felt implicit and never actually got made. It shows up three months later as an incident, an audit finding, or a re-architecture nobody budgeted for.
This is the list I run against every VCF 9.1 HLD before it leaves design review. Eighteen items, grouped by domain. Not exhaustive. But each one has cost somebody real time when it wasn't decided on purpose.

The diagram is the map for what follows. Same five layers this checklist is organised around, with each gap marked where it actually sits. Keep it open in a second tab while you read, or print it for the review.
Consumption Who Gets What, and How
GAP 1 The provider/consumer boundary gets drawn after tenants are already live
Org hierarchy, quota tiers, network isolation these are provider-admin calls, and they're expensive to change once real tenants depend on them. What usually happens: the platform team builds the capability, onboards the first eager application team, and only then works out what the tenancy model should have been. Decide the org/project structure, the quota approach (per-org or per-project, hard caps or soft with approval), and the network isolation boundary before your first real tenant, not after.
GAP 2 The catalog is built for new workloads, not the estate you actually have
Every self-service catalog ships full of templates for greenfield VMs and fresh Kubernetes clusters. Most of an estate's actual value sits in workloads that already exist. Decide up front whether and how existing production workloads get into the catalog App Stack Formation or something like it or just accept that self-service only ever covers new work while the bulk of the estate sits outside it.
GAP 3 Nobody has actually walked the consumer experience
Design reviews check the provider side org structure, network topology, policy far more often than they check the consumer side. If nobody on the design team has sat down and provisioned a workload through the catalog the way a tenant developer would, the design hasn't been tested from the side that actually determines whether people adopt it.
Operations The Day 2 Decisions Hiding in Day 0
GAP 4 Patch cadence policy treats every layer the same
VCF 9.1's patching model is really three different mechanisms with three different disruption profiles: management plane (no workload risk), control plane (Quick Patch, minutes), data plane (Live Patch where it applies). A change process that demands the same approval and the same maintenance window for a five-minute vCenter patch as for a full data-plane reboot cycle hasn't caught up with the platform. Write layer-specific change categories down before your first patch cycle. Don't let it default to whatever the last platform generation trained everyone to expect.
GAP 5 Live Patch eligibility gets assumed, not checked
Live Patch on ESX depends on the host being image-managed rather than baseline-managed, and on vSphere 8.x specifically, on TPM not blocking the fast path (that gap closed in 9.1). Estates that assume every host gets the quick patch path usually find out otherwise during their first Critical VMSA, not during design. Go through it cluster by cluster and write down which ones are actually on the fast path and which are still on the standard cycle. If any are on the slow path by accident rather than by decision, that's the one to fix first.
GAP 6 Nobody's written down who operates the consumption layer
Self-service moves operational ownership around catalog curation and policy sit with the provider team, but day-to-day workload operations shift to consumers who might have no infrastructure background at all. If the operating model doesn't say who's on the hook when a self-service database runs out of storage at 2am, the answer defaults to whoever happens to pick up the page. That's rarely the right person and never repeatable.
GAP 7 Release notes and BOM tracking has no owner
The Bill of Materials, patch releases, async releases these are living documents in the 9.x era, not something you check once at launch. Estates without a monthly review habit drift from documented state and nobody notices until an upgrade or an audit forces the issue. Give this an owner. It's cheap to do on a schedule and expensive to reconstruct after the fact.
Security and Compliance Where Assumptions Cost the Most
GAP 8 Compliance mapping is generic, not built for your estate
VCF 9.1's compliance tooling Advanced Cyber Compliance, STIG support, File Integrity Monitoring gives you the mechanism. It doesn't tell you which controls in APRA CPS 234/230, PCI DSS, the Essential Eight, or DORA map onto which platform settings for your estate specifically. That mapping is design work. It needs doing once per regulatory regime you're actually subject to, not assumed from a generic template someone used on a different customer.
GAP 9 Recovery is designed for the failure you can picture, not the one that happens
Fleet DR and Cyber Recovery (IRE) solve different problems infrastructure loss versus a compromised but technically-still-running estate. Estates often design and test one and assume it covers both. Recovery that depends entirely on the same provider or platform that just failed isn't really recovery. Check whether your architecture survives the failure of your primary platform or provider, not just a component within it.
GAP 10 Identity gets inherited, not designed
Estates coming from 5.x carry years of VIDM group structure and role mappings straight into Identity Broker without ever examining whether any of it should survive the move. The scripted VIDM-to-Identity-Broker migration is non-disruptive by design, which makes it easy to lift and shift years of group sprawl without deciding whether that's the access model you actually want. The migration window is the cheapest chance you'll get to clean up group-to-role mapping. Most estates don't use it that way.
GAP 11 DNS and certificate hygiene get treated as deployment mechanics
The strictly-lowercase FQDN requirement and certificate lifecycle planning both look like implementation detail until one blocks an integration or the other expires mid-production. Both are small enough to skip in a design review and expensive enough to earn one line each: confirm DNS records are checked for case sensitivity before staging, and confirm someone actually owns certificate issuance and renewal rather than assuming it's someone else's job.
GAP 12 The logging change doesn't make it into the SIEM config
Unencrypted syslog on port 514 is blocked in 9.1. The standard path now is TLS on 1514. Estates that haven't updated their SIEM and log-forwarding config to match find out about it as a silent logging gap, not a loud failure which is the worst possible way to lose visibility in exactly the system meant to catch problems.
Virtual Infrastructure What Storage and Networking Force You to Decide
GAP 13 vSAN feature adoption doesn't get checked against topology first
Global deduplication and a few other 9.1 storage efficiency features carry topology exclusions stretched clusters and 2-node configs are common gaps in current support. Estates that build the storage-efficiency business case before checking feature-to-topology compatibility end up re-justifying the TCO model after the fact. Check compatibility before the capacity plan gets built on top of it.
GAP 14 Memory tiering eligibility doesn't get audited against reservation settings
"Reserve all guest memory (All Locked)" permanently excludes a VM from memory tiering it's a hardware pin, not a scheduler preference while ordinary VM Memory Reservation doesn't. Templates inherited from pre-tiering estates often have All Locked set for reasons nobody remembers anymore. Audit for it before rolling out tiering. The gap between expected and actual consolidation savings is usually sitting in this one setting.
GAP 15 Connectivity model selection defaults instead of getting decided
Six connectivity models are supported, each with different operational and tenancy implications, and it's common to just default to whatever model the last deployment used rather than actually evaluate against current requirements self-service scope, NSX Federation, VPC consumption. Make the model selection an explicit, justified line in the HLD, not an inherited default nobody revisited.
Physical Infrastructure The BOM Decisions With Long Tails
GAP 16 Server certification gets checked once, at procurement, and never again
The Broadcom Compatibility Guide has three listing states supported, supported-with-vendor-confirmation, and unlisted and a server's status can be checked once at purchase and then never revisited. Firmware updates, component swaps, and BCG revisions can all move a configuration's status over its service life. Recheck BCG status as part of routine lifecycle review, not only when you're buying.
GAP 17 Multi-vendor adoption has no cluster design discipline behind it
Heterogeneous vSAN clusters are genuinely useful for phasing in new hardware, but mixing vendors without a deliberate policy produces exactly the operational cost multi-vendor strategies are supposed to avoid: divergent images, an undocumented add-on and firmware matrix, inconsistent support paths. Decide which clusters stay homogeneous (performance-critical, where consistency matters) and which can tolerate mixing (capacity, transition-phase) before procurement not as something that just happens because a vendor had a good quarter.
GAP 18 8.x carry-forward eligibility doesn't get checked before the refresh gets budgeted
A large share of vSphere 8.x-certified hardware carries forward into 9.x certification automatically, which means the refresh-or-keep decision should start with a BCG check, not an assumption that moving to 9.x means buying new servers. Estates that budget for a full refresh without checking carry-forward first are often solving a smaller problem than they think they have.

That layout isn't any specific customer's estate. It's the shape most of them take. Notice where the gaps cluster: the management domain carries the tenancy and patch-policy calls from the Consumption and Operations sections above, and the two workload domains show why "is this cluster Live Patch eligible" has a different answer depending on whether it's image-managed and homogeneous, or still carrying a baseline-managed, multi-vendor legacy. Same physical layout, two very different gap profiles.
How to Use This
Run this against your HLD before design sign-off, not after. For each item: either the document already has an answer, or it doesn't and "it doesn't" is a fine review outcome as long as it becomes an assigned action instead of a silent gap that turns into an incident later. The pattern across all eighteen is the same. None of these are missing platform capability. VCF 9.1 has an answer for every one of them. What's missing, when it's missing, is someone in the design process making the call on purpose instead of by default.
Sources
What each of these actually is, and which gap it backs, for anyone who wants to check the claim rather than take my word for it.
• Broadcom Modernizing Infrastructure: VMware Cloud Foundation 5.2.x to 9.1 Upgrade Guide
Broadcom's own walkthrough for moving a 5.2.x estate to 9.1. This is where the lowercase-FQDN requirement and the syslog port change (514 unencrypted blocked, 1514 TLS is standard) actually come from backs Gap 11 and Gap 12.
• Broadcom Faster Security Patching with Fewer Disruptions in VCF 9.1
The source for the three-layer patching model itself management plane, control plane, data plane, each with a different mechanism and disruption profile. Backs Gap 4.
• Broadcom Strengthen Zero Trust Platform Security and Resilience with VCF 9.1
Covers the TPM/Live Patch fix in 9.1 and the scripted VIDM-to-Identity-Broker migration. Backs Gap 5 (why TPM hosts used to miss the fast patch path) and Gap 10 (why the migration script doesn't force you to redesign access).
• Broadcom More Memory, Less Effort: Configuring Memory Tiering in VCF 9.1
A configuration-level post explaining exactly what "Reserve all guest memory (All Locked)" does versus ordinary VM Memory Reservation one pins pages to DRAM at the hardware level, the other doesn't. Backs Gap 14 directly.
• Broadcom More Capacity with VMware vSAN Compression and Global Deduplication in VCF 9.1
Explains the ZSTD compression change and where global dedup does and doesn't apply stretched clusters and 2-node topologies are currently excluded. Backs Gap 13.
• Broadcom VCF 9.0 Server Certification: Preserving Your Hardware Investment and Giving the Best ROI
Explains how the Broadcom Compatibility Guide actually works the three listing states (supported, supported-with-vendor-confirmation, unlisted) and how 8.x-certified hardware carries forward into 9.x. Backs Gap 16 and Gap 18.
• Broadcom Continuous Compliance, Integrated Cyber Recovery and Enhanced Platform Security for VCF 9.1
Introduces Fleet DR and Cyber Recovery (IRE) as separate blueprints solving separate failure modes infrastructure loss versus a compromised-but-running estate. Backs Gap 9.
• Broadcom Securing the Foundation: VMware Cloud Foundation 9.1 STIG Compliance
Covers Broadcom's STIG hardening tooling and how it maps to DoD-grade compliance baselines the same pattern applies to APRA, PCI DSS, and the Essential Eight, just with a different control set. Backs Gap 8.
• Broadcom The Future of Cloud Infrastructure: VCF Automation at Explore 2026
Covers App Stack Formation (turning an existing live VM into a catalog item) and the provider/consumer split in VCF Automation's operating model. Backs Gap 1 and Gap 2.
• Broadcom VCF Breakroom Chats Episode 92: Simplifying VCF 9.1 Install and Brownfield Import
A product-team walkthrough of the installer and the brownfield import workflow relevant background for Gap 1 and Gap 3, since import/onboarding is usually where the provider/consumer boundary first gets tested.
• Broadcom TechDocs VCF 9.1 Design (Modern Private Cloud design guide)
The official design guide, including the six supported connectivity models. Backs Gap 15.
• Broadcom TechDocs VCF 9.1 Release Notes
The living document itself Bill of Materials, patch releases, async releases. The thing Gap 7 says needs a named owner checking it monthly.
Related Field Guide Coverage
• Consumption: Part 61 (Adopting and Consuming VCF 9.1)
• Lifecycle and patching: the three-layer patching model, the security Express Patch series, Part 56 (baselines to images)
• Identity: Part 33 (Identity Broker and SSO), Part 40 (VIDM migration), the Identity Design Best Practices whitepaper piece
• Compliance: Part 46 (STIG), Part 38 (FIM), the regulatory compliance field guide
• Recovery: Part 35 (Cyber Recovery), the UniSuper case study
• Storage and memory: Part 42 (Memory Tiering What-If), the vSAN dedup/compression coverage
• Hardware: the ODM/OEM server economics piece, Part 47 (Supervisor load balancer)