Contents
1.3 Objectives and Requirements. 1
2.2 Prioritised Focus Areas. 4
3.6 Naming and Tagging Standards. 20
6.3 Role Based Access Control 50
6.5 Privileged Identity Management 54
8.2 People Process and Culture. 63
1 Introduction
1.1 Overview
RLC are conducting a top-down maturity assessment of the CUST Azure and DevOps implementation to ensure alignment with both the Microsoft Cloud Adoption Framework (CAF) and industry best practice and to assess the Landing Zones readiness to support future workload migrations.
RLC have compiled this document after working closely with key business, technical and security stakeholders to ensure current state, current challenges, and future aspirations of CUST implementation and maturity journey.
1.2 Document Purpose
The purpose of this document is to highlight the deltas between the current state of CUST’s Azure implementation and Microsoft’s Cloud Adoption Framework (CAF), whilst also ensuring industry standards and best practice.
This document will provide a current state view of each component within the scope of this document and when necessary, a proposed future state with a recommendation. This document will remain at a relatively high-level of detail to maintain the correct context of a maturity baseline. This is not a documented full CAF assessment and review but covering several key components.
1.3 Objectives and Requirements
The objective of the Azure Governance, Foundations, Cost Management & DevOps review and recommendations is determining current state maturity, along with series of recommendations to be considered and executed as a program of work to further implement best practices and strategies across the department.
Outcomes of the assessment aim to improve platform reliability, operational compliance and improve adoption for consuming teams.
1.4 Document Scope
The scope of this document is limited to the agreed scope that is outlined under the RLC-CUST-5368 statement of work.
1.4.1 In Scope
The following list details the items that are within the scope of this document.
It is important to note that whilst the CAF talks about people, process, and technology this assessment is target on the technology layer of the framework.
- Azure Foundational Design
- Azure Hierarchy
- Management Group Structure
- Subscription Structure and Decision Framework
- Resource Group Structure
- Policy
- Naming and Tagging Standards
- Cost Management
- Networking
- Virtual WAN
- Virtual Network
- Private Endpoints
- Firewalling
- Security
- Governance
- Policy
- Compliance
- Identity
- RBAC
- Conditional Access
- Privileged Identity Management
- Monitoring
- Logging
- DevOps
At the time of publication of this assessment access was limited to the following Azure tenants & subscriptions:
Tenants:
- CUST PTY LTD (CUSTptyltd.onmicrosoft.com)
1.4.2 Out of Scope
The following list details items that are specifically out of scope of this document:
- Creation of any resources
- Code quality and effectiveness assessment
- Operational Management
- Assessment of the On-Premises Environment
- Implementation of any recommendations
- Tenants & Subscriptions not listed in scope.
1.5 Target Audience
- RLC Solution Consultant – the requirements will be reviewed by the solution consultant to ensure that it is consistent with Microsoft’s Cloud Adoption Framework (CAF) and industry best practice.
- CUST key stakeholders – to help shape and formulate the assessment against key requirements gathered during workshops.
- CUST Compliance, Security and Cloud engineers and other technical personnel.
1.6 Document Conventions
| Recommendation |
| Recommendations will be captured in “Recommendation” boxes throughout the document. Each of these represents a recommendation from the team for future use and adoption of the solution but not delivered as part of the scope of this document. |
To consider the priority of the recommendations included in this section, the following simple structure rating of 1 to 5 (1 being Lowest and 5 being Highest) this provides an overview of each recommendation.
- Benefit:
- 1 = Low Business / Operational Benefit
- 5 = High Perceived Business / Operational Benefit (or Timesaving)
- Effort:
- 1 = 5 Days
- 5 = Effort will take 3 months +
- Complexity:
- 1 = Simple change no dependencies
- 5 = Highly complex change with multiple dependencies.
| Recommendation | Benefit | Effort | Complexity |
| Network: Implement Azure VWAN | 5 | 3 | 4 |
2 Recommendations
Recommendations in this section are consolidated from the detailed review sections that follow.
Specific identification has been made for uplifts that need to occur to improve the Azure platform at CUST and the Effort verses Benefit for these changes.
As an example, a reimplementation to the service is a redeployment, a refactoring to a service is a change/reconfiguration to an existing service that can be re-integrated with the existing implementation without wholesale change.
2.1 Summary
Our assessment of the Azure environment confirms it provides a suitable foundation for CUST’s current and planned workload requirements. Through our evaluation, we examined several core technical domains to assess the platform’s alignment with Microsoft landing zone best practices and its capability to support future workload migrations.
The environment demonstrates fundamental controls and architectural elements that enable secure and scalable operations. Our review covered Identity and Access Management, Resource Organisation and Standards, Network Topology and Connectivity, Management, Governance, Platform Automation and DevOps, and Security.
Throughout this report, we detail our findings across these domains and identify specific areas where the environment can be enhanced to increase operational maturity and better align with Microsoft best practices. These opportunities for improvement range from tactical adjustments to strategic initiatives that will strengthen the platform’s capabilities.
The subsequent sections provide detailed analysis of each domain, including current state assessment, alignment with best practices, and specific recommendations for improvement. These recommendations are presented with consideration for both immediate operational needs and long-term strategic objectives.
Our assessment indicates that while the platform is ready for use, implementing the recommended improvements will further enhance its security, scalability, and operational efficiency. There is, however, a key callout noting that there are no Architectural artefacts that accurately describe the built platform and landing zone patterns.
2.2 Prioritised Focus Areas
The following pages of this document present detailed observations and recommendations about CUST’s Azure cloud environment. From all those recommendations, the top focus areas are listed in the table below.
| Priority | Section | Focus Area | Comment |
| 1 | Azure Hierarchy | Resource Organisation | · Simplify the Azure scaffold structure by grouping management groups more efficiently e.g. Business units or Functions. · Configure a ‘default’ management group for new subscriptions so that they are not created under the Root management group. · Align the enterprise agreement department and accounts to reflect the relevant business structure. · Enable cost management by leveraging budgets and budget alerts to track costs. |
| 2 | Management | Logging | · Centralise logging (and enable consistently). · Consider segregating Operational logging data from Security logging data to save on cost. · Establish Logging patterns for all services (PaaS and IaaS), document and deploy consistently via Azure Policy. · Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data. · Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards. · Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure. |
| 3 | Operations | Monitoring | · Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone. · Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example). · Configure Service and Resource health alerts and notify central platform and operational teams when issues occur. · Define and implement Monitoring and alerting on all PaaS Services. · Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications. · Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns. |
| 2 | Network | Topology, Connectivity and Security | · Review the current Network architecture and consider uplifting to Azure Virtual WAN to ensure a future proof scalable model and simplify routing. · Ensure NSG’s are applied to all subnet and enforced with Azure Policy to support granular rulesets. · Consider the use of Security Admin Rules to enforce mandatory network ACL’s. · Secure Core network from standard access to improve Security. · Enable connectivity monitors with Alerts on key network resources. · Enforce the use of Private Endpoints on supported services using Azure policy. · Periodically review traffic analytics to ensure network flows and traffic are as expected and manage the exceptions. · Leverage a Hierarchical approach to applying Azure firewall policies (Use A Global policy for Core rules). · Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny. · Configure Azure firewall IDS in Deny mode. · Consider using IP groups to reduce IP table rules usage. |
| 3 | Foundations | Architecture and Standards | · Align design and as-built documentation to the current platform architecture and state. · Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework. · Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables). · Create solutions module repository of common reference Terraform modules for deploying Azure resources. |
| 4 | Governance | Governance and Cost Management | · Adopt a Policy driven approach to enforcing guardrails for Azure resources. · Leverage Azure policy to govern which Azure services are allowed to be deployed by Application teams. · Create Azure resource locks to prevent accidental deletion of critical resources, these can be managed as a stage in the DevOps deployment pipelines. · Create a Show Back / Charge Back model for cost management (enable a FinOps Approach) · Create a process to manage (and delete) Orphaned Disks and NICs · Ensure cost definitions are built into the inception of a services creation. · Action Azure Advisor recommendations where it makes sense. · Action Azure Advisor insights. · Consolidate Logging then apply Capacity on these spaces |
| 4 | Identity and Access Management | Role Based Access and Authentication | · Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments. · Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones. · Consider establishing a cloud-based remote-access solution for Azure Administration. · Re define PIM to audit and validate that PIM escalation is tightly controlled. · Conduct PIM access reviews periodically. |
| 6 | Automation and DevOps | Architecture and Azure DevOps | · Define the DevOps standards and principles your teams must follow. · Align the DevOps strategy with the business objectives. · Develop and define key metrics for the DevOps team to improve deployment related aspects of the platform. · Evaluate the effectiveness of the current DevOps practices and toolchain in relation to the Enterprise Landing Zone. · Determine and document how the Landing Zone team should collaborate with other teams within the organisation to effectively capture requirements and iterate on design and the platform. · Develop a centralised reusable set of version-controlled blueprint or modules for high-use archetypes. |
3 Azure Hierarchy
3.1 Overview
| Figure 3‑1: Azure Hierarchy |
Azure provides four levels of management scope: those being Management Groups, Subscriptions, Resource Groups and Resources. Ensuring there is a defined structure to this management scopes is critical to managing and securing your cloud environments whilst also ensuring maximum visibility into the use and cost of your resources in Azure.
3.2 Management Groups
3.2.1 Definition
Management Groups are containers that help you manage access, policy, and compliance across multiple subscriptions. All subscriptions in a management group automatically inherit the conditions applied to the management group.
3.2.2 Current State
CUST are currently using management Groups with the structure outlined in the below image.
There is clear delineation between the Platform and Landing Zone management groups. New Landing Zones will be built under the ‘mg-application’ management group, with the depicted structure of a management group and subscription per application. The ‘mg-platform’ management group contains sub management groups for core services which are shared by all onboarding Landing Zones.
There are distinct management groups for security resources, enabling clear segregation and sandbox environments, which will be leveraged for PoC (Proof of Concept) activities.
Figure 3‑2 – Current Management Group Structure
3.2.3 Cloud Adoption Framework (CAF)
Microsoft’s Cloud Adoption Framework provides several examples of grouping subscriptions within management groups as well as proposing the use of a “Mixed” strategy for complex organisations which combine one or two of the primary strategy examples. Below is a table of these several examples:
| Strategy Name | Description |
| Workload Separation | Subscriptions are categorised by their environment type (Production or Non-Production/Pre-Production) and placed into management groups that reflect this. |
| Application Category | Subscriptions are again categorised by environment type although there are usually more environment types in this approach. (Production, Development, QA etc.). There are also subscriptions created for specific workloads, such as Mission Critical or Protected Data. |
| Functional | Subscriptions are categorised by functional streams such as finance, sales or IT and are placed into management groups to reflect this. |
| Business Unit | Subscriptions are logically groups according to the business units they belong to. Business units can be defined by profit and loss category, division, profit centre or business structures. |
| Geographic | Subscriptions are grouped within management groups based on their geographic distribution, with a management group representing each geographic location. |
| Mixed | Mixed strategies comprise any two (or more) of the above 5 mentioned strategies together to define a more complex but granular hierarchy. |
3.2.4 Recommendation and Assessment
CUST are currently using Management Groups and have a good implementation mostly aligned with best-practice. Although CUST doesn’t have a formalised Cloud Operating Model, the current setup is most closely aligned to the Centralised operating model of the Azure tenant by the Infrastructure and Security chapter. The current management structure supports this, but it could also support scaling to an Enterprise operating model in line with CUST’s cloud adoption growth ambitions. The Cloud Operating model should be reviewed as part CUST’s Cloud strategy. More on that can be found here.
The Applications Management Group should be renamed to ‘mg-landing-zones’ to reflect the purpose and remove the application specific Management Groups. If required, introduce another Management Group layer under ‘Landing Zones’ with an appropriate grouping e.g. Business Unit.
In the future, the platform team may consider an Archive management group to store offboarded landing zones where there is a requirement to maintain the state or data for a period after decommissioning.
Figure 3‑3: Recommended Management Group Structure
Figure 3-3 illustrates the recommended management group structure.
| Management Group Recommendation |
– Decommission the management group per application structure and introduce a landing zone centric approach. – If an additional layer is required, introduce it as per the ‘Optional’ area depicted above. – Formally document the cloud operating model. – In the future, introduce the Archive/Decommissioned management group so that landing zones that offboarded have a holding zone in case of any retention purposes. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Decommission the management group per application structure and introduce a landing zone centric approach. | 4 | 2 | 1 | REFACTOR |
| If an additional layer is required, introduce it as per the ‘Optional’ area depicted above. | 4 | 2 | 1 | REFACTOR |
| Formally document the cloud operating model. | 5 | 3 | 1 | REFACTOR |
| In the future, introduce the Archive/Decommissioned management group so that landing zones that offboarded have a holding zone in case of any retention purposes. | 3 | 1 | 1 | REFACTOR |
3.3 Subscriptions
3.3.1 Definition
Subscriptions logically associates user accounts and the resources that were created by those user accounts. Each subscription is subject to limits and quotas on the number of resources that can be created or used. Organisations can use subscriptions to manage costs and the resources that are created by users, teams, or projects.
3.3.2 Current State
CUST have created implemented the following subscriptions based on application environments largely based on a clear structured naming standard.
Subscription tagging has been implemented and follows a consistent and concise structure, with inheritance which propagates tags onto resource groups.
| Subscription Name | Resource Group Locations | Secure Score |
| CUSTNP-Azure-Test | Australia East | – |
| CUSTNP-CAMMS | None | – |
| CUSTNP-Connectivity | Southeast Asia Australia East | – |
| CUSTNP-Deltek | None | – |
| CUSTNP-Integration | Australia East | – |
| CUSTNP-Jira | None | – |
| CUSTNP-Management | Southeast Asia Australia East | 36% |
| CUSTNP-Predict | None | – |
| CUSTNP-Primavera | None | – |
| CUSTNP-Sandpit | East US Southeast Asia Australia East Australia Southeast | 62% |
| CUSTNP-Security | Australia East | 20% |
| CUSTPD-CAMMS | None | – |
| CUSTPD-Connectivity | Australia East Australia Southeast | 84% |
| CUSTPD-Deltek | None | – |
| CUSTPD-Fabric | Australia East | 20% |
| CUSTPD-HybridARC | Australia East | 20% |
| CUSTPD-Identity | Australia East | – |
| CUSTPD-Integration | Australia East Australia Southeast | – |
| CUSTPD-Internet | Australia East | 39% |
| CUSTPD-Jira | None | – |
| CUSTPD-Management | Southeast Asia Australia East Australia Southeast | 53% |
| CUSTPD-NetworkTeam | Australia East | 20% |
| CUSTPD-Predict | None | – |
| CUSTPD-Primavera | None | – |
| CUSTPD-Security | Australia East Australia Southeast | 89% |
Visual Studio Enterprise Subscription 3b5cbcbf-c2ec-449e-8a22-04a292a4e2d2 | None | – |
Visual Studio Enterprise Subscription 8a502b0d-c9e0-427e-b94b-936e5de89cb2 | None | – |
In the Subscriptions policy, the settings are currently set to their defaults. These policies govern what level of authorization is required to move subscriptions to and from Azure Active Directory tenants.
With the current default values, it presents the risk that a non-authorized user account could move a subscription to a non-CUST managed tenant (thus removing all management controls).
Figure 3‑4 – Subscription Policies
3.3.3 Cloud Adoption Framework (CAF)
Azure Cloud Adoption Framework recommends organising subscriptions based on both scale and functional requirements, serving as fundamental containers for resources with distinct management, billing, and access control boundaries. Subscriptions should be structured to support organisational requirements for compliance, security, and operational efficiency, while considering service limits and scale constraints. The framework advocates for a hierarchical management approach using management groups for policy inheritance and governance, with subscriptions organised to separate different environments (production, development, test), business units, or workload types. This organisation should align with the enterprise’s billing requirements, administrative model, and security policies while maintaining clear isolation boundaries where needed. Resource management, policy enforcement, and access controls are implemented at the subscription level, making it crucial to plan subscription architecture that accommodates both current needs and future growth while enabling effective cost management and operational oversight.
3.3.4 Recommendation and Assessment
With a new management group structure simplifying logical boundaries, we are freed up to utilise more granular subscription based on one logical boundary (environment) and take advantage of the inheritance that is offered by the hierarchy. This will reduce the complexity and management overhead of managing subscriptions.
Due to the desire to leverage a landing zone per application in the future, it is important to leverage Azure native cost management tools and implement a detailed cost management plan.
3.4 Resource Groups
3.4.1 Definition
Resource Groups act as logical containers into which Azure resources can be deployed and managed. They act as one of the foundational boundaries for cost, security, and compliance within Azure. Every resource in Azure must belong to a resource group.
3.4.2 Current State
82 Resource Groups have been observed in Azure at the time of publishing this assessment.
There is some inconsistency regarding the naming, capitalisation and logical boundary used to group resources, largely due to the Microsoft default resource groups deployed with some of their solutions e.g. Log Analytics Workspace.
3.4.3 Cloud Adoption Framework (CAF)
Resource Groups are the last logical grouping when it comes to the Azure Hierarchy. Based on the decisions that are made regarding the Management Group and Subscription structures, Resource Groups should be used to group resources that share the same lifecycle or security boundary.
3.4.4 Recommendation and Assessment
Resource Groups should be named relatively simply concentrating on the context of the lifecycle it is grouping, as well as the environment that it belongs to. In the current Azure implementation, it is unclear what the majority of Resource Groups are grouping.
| Resource Group Recommendation |
– Investigate and document the lifecycle of Resource Groups. – Remove empty resource groups . – Review resource groups that have been autogenerated and do not conform to the CUST Naming Standards. – Enforce naming standard via a naming module and Azure Policy. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Document the lifecycle of Resource Groups. | 3 | 2 | 2 | REFACTOR |
| Remove empty Resource Groups. | 1 | 1 | 1 | REFACTOR |
| Review Resource Groups that do not conform to the CUST naming standards (including autogenerated) | 2 | 1 | 1 | REFACTOR |
| Review and enforce Naming Standard via naming module and Azure Policy | 4 | 3 | 3 | REFACTOR |
3.5 Policy
3.5.1 Definition
Azure Policy is a service that enables the creation, assignment, and management of policies. Policies are used to enforce rules on your resources to ensure they remain complaint to the standards defined by your corporate standards. Azure Policy scans across your resources to identify resources that are not compliant with the implemented policies.
Policies can be assigned at varying scopes, from management groups to resource groups. These offer a method for enabling organization standards from the root management group, with geographic standards applied at their respective management groups.
3.5.2 Current State
Azure Policy is used extensively at CUST with 7262 policies currently in use (mainly compliance Initiatives) and an overall Resource compliance of 90%.
Other policies to be implemented to control the use of:
- VM SKUs
- Permitted resource types.
- Permitted Region deployment
There are several remediation tasks which are pending form the remediation blade in Azure policy
3.5.3 Cloud Adoption Framework (CAF)
The CAF outlines Enterprise Governance as a team sport, requiring all parties to be involved in the process. There is a large amount of effort that comes with translating on-premises based IT policies to the cloud, often requiring the Platform teams and Security to collaborate to translate the policies accurately with the correct language and scope.
Azure Policy within the CAF focuses on four main areas:
- Cost Management
- Identity Baseline
- Security Baseline
- Resource Consistency
Enabling the use of additional policies is the primary step to ensuring compliance with the CAF, some of the policies recommended:
- Enable Specific VM SKUs.
- CIS Audit Initiative
- Allowed Resource Types (Used to Deny services that have not been CUST ‘endorsed’)
3.5.4 Recommendation and Assessment
Enabling the use of the CIS Audit Initiative will give a grant insight into alignment against the CIS Azure Platform Benchmark. The Policies within the CIS Benchmark align to a great deal of the policies that CUST have currently implemented but are constantly updated by Microsoft with new policies that further align to the CIS benchmark.
Azure Policy can be used to assess compliance of configuration to security standards, defining the expected security standards for the top 10 most used Azure resources and enabling policies to report on the level of compliance would give great insight to the maturity of the Azure environment against the expected documented policies.
| Azure Policy Recommendation |
– Create an Azure Policy Governance Framework and enable a team to review, remediate, communicate, and action policies. – Define a policy of allowed resource types that have been ‘whitelisted’ for use within the CUST Azure environment. – Create a process for onboarding new services and defining appropriate policies to apply to these. – Define the use of Azure Policy for the Organisation, who owns and communicates these and how they are life cycled. – Consider assigning the Microsoft Defender policies at the Management Group level and unassign them at the subscription level to avoid repeated assignments. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Create an Azure Policy Governance Framework and enable a team to review, remediate, communicate, and action policies. | 5 | 3 | 3 | REFACTOR |
| Define a policy of allowed resource types that have been ‘whitelisted’ for use within the CUST Azure environment. | 5 | 2 | 2 | REFACTOR |
| Create a process for onboarding new services and defining appropriate policies to apply to these. | 5 | 3 | 3 | REFACTOR |
| Consider assigning the Microsoft Defender policies at the Management Group level and unassign them at the subscription level to avoid repeated assignments. | 4 | 1 | 1 | REFACTOR |
3.6 Naming and Tagging Standards
3.6.1 Definition
A well-defined naming and tagging standard are a key component to ensuring the ability to quickly locate and manage resources for operational purposes, as well as help to associate usage costs with specific resources and enable a chargeback or show-back model if required. It is important to note that within Azure resource types have different scopes that define the level at which the name must be unique.
3.6.2 Current State
Naming Standards
CUST have implemented a well-structured naming standard based on the Azure resource type, hosting subscription, environment, region and purpose. The Naming standard has been consistently applied except on resources and resource groups that are auto provisioned such as Log Analytics.
Tagging Standards
CUST have implemented a standardised tagging model that are consistently deployed using Azure Policy. The tags that are deployed are a minimal set to identify ownership and could be uplifted to enhance the context of the resource and include operational metadata scenarios that would be required and useful.
It was noted that there are several resources have a tag, but there is no value rendering the tag invalid.
3.6.3 Cloud Adoption Framework (CAF)
Naming Standards
A good naming standard helps to identify resources, be that in cost reporting or in automation scripts. There are two main components that a naming strategy should include, those being Business and Operational.
- The business-related component to a naming standard should include the organisational information that is required to identity the teams responsible for the resource.
- The operational side should ensure that names include relevant information that IT teams need, to ensure that the resource can be identified efficiently.
As previously mentioned, CUST have a well-structured naming standard with adequate contextual components to assist in identification of the resource.
Tagging Standards
Tags are essentially a quick way to identify resources via metadata. Tagging is usually very specific to each organisation but the CAF covers a standard examples list of tags that can be used as a template to decide what tags are relevant to your preferred level of scope.
3.6.4 Recommendation and Assessment
The naming convention at CUST should be extended upon to meet the CAF recommendations. This will help with resource scale, identification, and untimely management. Particularly in an environment such has CUST a comprehensive naming standard will assist with management and governance.
There should also be specific tags created for functionality such as automated shut down and startup of virtual machines based on a schedule, but these are specific to virtual machine resources and are an additional capability.
| Naming and Tagging Recommendations |
– Review the list of required tags and uplift the tagging standards if required to include additional business and operational context. – Re-develop the current tagging policy to enforce the mandatory set of tags, especially on resources. – Remediate and update missing tag values. – Define and Communicate the Tagging Schema to all Cloud Consumers. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Review the list of required tags and uplift the tagging standards if required to include additional business and operational context. | 4 | 2 | 2 | REFACTOR |
| Re-develop the current tagging policy to enforce the mandatory set of tags, especially on resources. | 4 | 1 | 1 | REFACTOR |
| Remediate and update missing tag values. | 5 | 1 | 1 | REFACTOR |
| Define and Communicate the Tagging Schema to all Cloud Consumers. | 3 | 1 | 1 | REFACTOR |
3.7 Cost Management
3.7.1 Definition
Cost Control or Management, much like governance and similar management constructs, depend on a well-managed environment. It is important that the desired level of cost management is factored into the overall design of the Azure Hierarchy. Azure Cost Management gives you the tools to plan for, analyse and reduce your spending to maximise your investment in cloud.
3.7.2 Current State
At present, most of the Azure environment is centrally funded by IT with no show/charge back.
This could result in a cost-blowout as Teams have no ownership for costs and when it comes to provisioning their resources when they start consuming the platform.
In conjunction with a considered tagging schema, cost-show back or cost-chargeback can assist in cost-management.
Motivating resource consumers to:
- right-size,
- right-time, and
- deprovision unused resources.
Creating a culture of Cost Management for all Cloud resource ensures that CUST get value out of the services they build, as a stretch target a cost per product should be aimed for, this can them be linked to a value of the product to the business.
The use of budgets, cost alerts & reserved instances have not been implemented.
As CUST progress to consuming this platform and migrating workloads into Azure, it is important to consider the following cost management strategies from the start to avoid unwanted cost.
- Right Sizing: Enabling Azure monitor for VM’s to collect usage and thus gain insights into right sizing (less than 5% CPU use).
- Right Timing: extended use of power cycling in the environment would assist inreducing costs (i.e. turn off when not needed servers).
- Define ingestion patterns for Logs, presently LAW is centralised to the Sentinel enabled workspace, a potential shift to an operational LAW may be beneficial to support cost management.
- Ensuring the use of Hybrid licensing for SQL and Windows licenses.
Multiple Orphaned Disks and Network cards shows that resources being deleted are not cleanly done, adding additional costs.
3.7.3 Cloud Adoption Framework (CAF)
Cost Management within the CAF is a large point of discussion as most customers as they move to the cloud are extremely concerned about bill shock or unforeseen costs. Cost Management and Cost Tracking is intended to be accomplished within Azure with the use of Azure Cost Management, Azure Advisor, Tagging, Budgets and Alerting. It is also important to ensure that cost management is taken into consideration from a subscription and management group design point of view as these scopes can enable quick visibility, but also ensure that ability apply the correct permissions for cost management at a higher level.
3.7.4 Recommendation and Assessment
To ensure visibility of cost across Azure, default budgets and alerts should be configured within each subscription or management group to alert budget owners about potential unexpected spend or growth within their environments.
| Cost Management Recommendation |
– Configure the Enterprise portal to accurately reflect the organisation structure or cost using departments and accounts. – Create a Show Back / Charge Back model for cost management (enable a FinOps Approach). – Remove un needed Disks that are not attached to save money (example workbook). – Enable diagnostics on VM’s (and PaaS Services) to allow for right sizing recommendations. – Enable sensible default budgets and allow teams to create budgets so that costs can be managed and support review cycles. – Create a process to manage and report (and deleted) Orphaned Disks and Nics. – Action Azure Advisor insights. – Ensure cost definitions are built in to the inception of a services creation . |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Configure the Enterprise portal to accurately reflect the organisation structure or cost using departments and accounts. | 5 | 2 | 2 | REFACTOR |
Create a Show Back / Charge Back model for cost management (enable a FinOps Approach).
| 4 | 2 | 2 | REFACTOR |
| Remove un needed Disks that are not attached to save money (example workbook). | 3 | 1 | 1 | REFACTOR |
| Enable diagnostics on VM’s (and PaaS Services) to allow for right sizing recommendations. | 4 | 1 | 1 | REFACTOR |
| Enable sensible default budgets and allow teams to create budgets so that costs can be managed and support review cycles. | 5 | 1 | 1 | REFACTOR |
Action Azure Advisor insights.
| 5 | 2 | 1 | REFACTOR |
Ensure cost definitions are built into the inception of a services creation.
| 4 | 2 | 2 | REFACTOR |
4 Network
4.1 Overview
Azure has evolved over time to have numerous ways in which networks can be configured for various enterprises. Currently the CAF has two main network topologies that are recommended for implementation, those being, Hub and Spoke and Virtual WAN. As of the start of this month (July 2020), the recommended approach for large enterprise is to leverage the Virtual WAN capability instead of the traditional Hub and Spoke Methodology.
4.2 Virtual WAN
4.2.1 Definition
Azure Virtual WAN is a Microsoft-managed solution where end-to-end global transit connectivity is provided by default. Virtual WAN hubs eliminate the need to manually configure network connectivity. For example, you don’t need to set up or managed user-defined routing (UDR) or network virtual appliances (NVAs) to enable global transit connectivity.
Virtual WAN greatly simplifies the end-to-end network connectivity in Azure and cross-premises, by creating a hub and spoke network architecture that spans multiple Azure regions and on-premises locations.
4.2.2 Current State
CUST have implemented a hub and spoke network design managed by a combined effort of the Network team and the Cloud and DevOps team.
There is a single 50Mbps Express route circuit provided by MegaPort which provides connectivity to the CUST core network. To encrypt Express route, there are 2 VPN’s
- Australia East – Azure VPN Gateway over express route private peering
- Australia Southeast – Palo Alto VPN over express route in Australia East.
This network configuration is CUST current approved network topology for Azure connectivity.
Figure 4: Existing Network Topology
4.2.3 Cloud Adoption Framework (CAF)
A network topology based on Azure Virtual WAN is the preferred and recommended enterprise-scale approach for large-scale multi-region deployments where your organisation needs to connect your global locations to both Azure and on-premises locations. A Virtual WAN topology should also be used whenever your organisation intends to use software-defined WAN (SD-WAN) that are integrated with Azure. As a Microsoft-managed service it also reduces the complexity of your network and helps to modernize the landscape.
4.2.4 Recommendation and Assessment
CUST have implemented a well-designed and managed network topology for extending the core network to Azure. This design aligns well to the CAF recommendation for a traditional hub & spoke network topology.
Virtual WAN is a topology for CUST presents an opportunity to utilise the best of both words in terms of unlocking the network landscape within Azure, but also maximizing the investment that has already been made into the existing network.
Microsoft have a guide specifically for the migration of existing global Azure and on-premises footprints to Azure Virtual Hub.
Changing the existing network topology unlocks an opportunity to move away from using networks as a primary security boundary as we did in the previous on-premises worlds, enabling a move towards Zero-Trust networking.
Figure 5: Example of Azure Virtual WAN Topology
The above image depicts a high-level view of a simplistic target state for Azure Virtual WAN.
Benefits:
- Move away from Network as the primary security boundary (move towards zero – trust)
- Simplification – Routing intent negates the use of UDR and enables Plug and Play Connectivity for Branches.
- Alignment to Microsoft Recommendations for large enterprise connected environments.
| Networking Topology Recommendation |
– Investigate migrating to a Virtual WAN based Topology to unlock network capability within Azure. – Integrate the Azure firewall into Virtual WAN to create a Secured hub core routing capability. – Consider using different Express route circuits from other peering locations to ensure redundancy. – Configure alerting on Core network services including Express route circuits. – Enable connection monitor to monitor connectivity across Express route and VPN. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Investigate migrating to a Virtual WAN based Topology to unlock network capability within Azure. | 5 | 5 | 4 | REIMPLEMENT |
| Integrate the Azure firewall into Virtual WAN to create a Secured hub core routing capability. | 5 | 2 | 3 | REFACTOR |
| Consider using different Express route circuits from other peering locations to ensure redundancy. | 5 | 3 | 3 | REFACTOR |
| Configure alerting on Core network services including Express route circuits. | 5 | 1 | 1 | REFACTOR |
Enable connection monitor to monitor connectivity across Express route and VPN.
| 5 | 1 | 1 | REFACTOR |
4.3 Virtual Networks
4.3.1 Definition
Azure Virtual Networks (VNet) are the fundamental building block for your private network in Azure. VNet enables many types of Azure resources, such as Azure Virtual Machines (VM), to securely communicate with each other, the internet, and on-premises networks. VNet is like a traditional network that you would operate in your own datacentre but brings with its additional benefits of Azure’s infrastructure such as scale, availability, and isolation.
4.3.2 Current State
CUST currently have 21 observable Virtual Networks deployed within Azure.
There is a single Express Route 50mbps throughput and no redundant links.
The VNETS are linked to a hub and spoke model and are not separated by Environment and leverage a common firewall and VNET gateway for traffic flow.
Azure Virtual Network Manager and network groups are used to configure the connectivity and peering from spoke VNET’s to the Hub/Core Azure firewall VNET.
Great use of Virtual Network flow logs in the environment, this has been consistently deployed.
Within the 21 observed VNETs, it was noted that several subnets did not have Network Security Groups (NSGs) associated. Additionally, several subnets lacked route tables. Both conditions require review, consideration, and potential remediation.
List of Subnet without an NSG associated:
| VNET Name | Subscription ID | Subnet Name |
| vn-sdp-ae-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | AzureBastionSubnet |
| vn-sdp-ae-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-dsclab |
| vn-pmv-ae-prd-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-app |
| vn-pmv-ae-prd-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-sqldb |
| vn-internet-ae-prd-eslz-001 | 457f2bde-b26b-4950-8c9a-1aab67bf63eb | sn-management |
| vn-internet-ae-prd-eslz-001 | 457f2bde-b26b-4950-8c9a-1aab67bf63eb | sn-pe |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-inb |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-ha |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-out |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | AzureBastionSubnet |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-ha |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-inb |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | AzureBastionSubnet |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-out |
List of Subnets without route tables associated:
| VNET Name | Subscription ID | Subnet Name |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-out |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-inb |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | AzureBastionSubnet |
| vn-con-ae-prd-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-ha |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-ha |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | AzureBastionSubnet |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-out |
| vn-con-as-drp-eslz-001 | e7d313ee-9482-4278-b7aa-adb964578694 | sn-dns-inb |
| vn-internet-ae-prd-eslz-001 | 457f2bde-b26b-4950-8c9a-1aab67bf63eb | sn-management |
| vn-internet-ae-prd-eslz-001 | 457f2bde-b26b-4950-8c9a-1aab67bf63eb | sn-pe |
| vn-pmv-ae-prd-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-sqldb |
| vn-pmv-ae-prd-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-app |
| vn-sdp-ae-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | sn-dsclab |
| vn-sdp-ae-eslz-001 | f8fc7c8d-53ca-454e-b4d8-45671e7c9b9f | AzureBastionSubnet |
4.3.3 Cloud Adoption Framework (CAF)
The CAF when aligning to recommendations of Enterprise Scale, recommends the use of capabilities such as:
- Zero Trust Networking.
- Network Security Groups use Azure Service Tags rather than IP Addresses.
- Delegate subnet creation to the Landing Zone / App Resource Group, Owner.
Virtual Networks are the spokes within our Hub and Spoke network topology, even when utilising the Virtual WAN Capabilities.
4.3.4 Recommendation and Assessment
Having an overall Hub-Spoke architecture on the Azure Virtual networks already in is a good starting point. There are some improvements that can be made to bring it into more alignment with the Cloud Adoption Framework for an enterprise environment. These will secure resources and Network flows plus allow for more observability of traffic.
| Networking Recommendation |
– Review the deployment of NSGs to all Subnets and remediate where necessary. – Review the deployment of route tables to all Subnets and remediate where necessary. – Implement a policy to enforce NSG’s and Route tables on all subnets. – Uplift Azure Virtual Network Manager and implement Security Admin rules for default platform wide rules. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Review the deployment of NSGs to all Subnets and remediate where necessary. | 5 | 2 | 2 | REFACTOR |
| Review the deployment of route tables to all Subnets and remediate where necessary. | 5 | 2 | 3 | REFACTOR |
| Implement a policy to enforce NSG’s and Route tables on all subnets. | 5 | 3 | 3 | REFACTOR |
| Uplift Azure Virtual Network Manager and implement Security Admin rules for default platform wide rules. | 5 | 3 | 2 | REFACTOR |
4.4 Private Endpoints
4.4.1 Definition
Azure Private Endpoints is a network interface that connects you privately and securely to a service powered by Azure Private Link. Private Endpoints uses a private IP address from within your virtual network, effectively bringing the service into your virtual network. This can be used by Azure services such as Azure Storage, Azure Cosmos DB, SQL etc.
4.4.2 Current State
Private Endpoints are currently utilised across the CUST azure network, these are mostly made up of connections to blob endpoints and key vaults, although there is a wide range of other services included.
4.4.3 Cloud Adoption Framework (CAF)
The CAF currently only mentions Service Endpoints, however the same justifications or benefits for the use of Service Endpoints also applies to Private Endpoints.
By utilising Private Endpoints, CUST can effectively pull Azure PaaS resources within our virtual network’s private IP range, when enables secure and private connection to Azure resources. Providing benefits such as:
- Improved Security for Azure Resources
- Optimal Routing for Azure Resource traffic from within the Virtual Network
- Minimal Management Overhead
4.4.4 Recommendation and Assessment
Azure Private Endpoints should be enabled where a private facing PaaS resource is required, if there is no Private Endpoint available for that PaaS resource Service Endpoints should be utilised.
| Private Endpoint Recommendation |
– Create a Security Pattern to always use Private Endpoints where available and enforce using Azure policy. – Use Service Endpoints when Private Endpoints are not available with a routable Service Endpoint subnet, managed by Azure Firewall and Service Endpoint policies attached. – Use Azure policy to create and lifecycle manage DNS records associated with Private endpoints. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Create a Security Pattern to always use Private Endpoints where available and enforce using Azure policy. | 4 | 2 | 2 | REFACTOR |
| Use Azure policy to create and lifecycle manage DNS records associated with Private endpoints. | 5 | 2 | 2 | REFACTOR |
4.5 Azure Firewall
4.5.1 Definition
Azure Firewall is a managed, cloud-based network security service that protects your Azure Virtual Network resources. It is a fully stateful firewall as a service with built-in high availability and unrestricted cloud scalability.
4.5.2 Current State
CUST physical network connections are configured via express rout circuits with an Azure VPN overlay in Australia East and Palo Alto VPN for Australia Southeast using the same Express route circuit.
All traffic is routed through the Azure Firewalls deployed in the Hub VNET and routing enforced using a route table where it exists.
Whilst this is a legacy Hub and Spoke pattern, it is still effective but could certainly be optimised using the Virtual WAN Secure Hubs as mentioned earlier in the document.
4.5.3 Cloud Adoption Framework (CAF)
The CAF largely refers to the use of Azure Firewall via design recommendations:
- Azure Firewall should be deployed in the Virtual WAN Hubs for east-west and/or north-south traffic protection / filtering.
- Within a Hub and Spoke topology Azure Firewall should be deployed within the Central HUB for east-west and/or north-south traffic protection / filtering.
- Create global Azure Firewall Policy to govern security posture across the global network environment and assign it to tall Azure Firewall instances.
- Enable the use of Azure Service Tags where possible instead of using dedicated IP Addresses.
4.5.4 Recommendation and Assessment
CUST have implemented a well-designed and managed network topology for extending the on-premises network to Azure. This design aligns well to the CAF recommendation for a traditional hub & spoke network topology but could be optimised to leverage Azure Virtual WAN.
| Azure Firewall Recommendation |
– Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs. – Uplift the firewall policy and create a Hierarchy, set a Global policy to all existing firewalls that contains ‘must have’ rules and use local policies for individual and localised rules. – Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs. – Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny. – Configure Azure firewall IDS in Deny mode. – Consider using IP groups to reduce IP table rules usage. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs. | 5 | 4 | 3 | REIMPLEMENT |
| Uplift the firewall policy and create a Hierarchy, set a Global policy to all existing firewalls that contains ‘must have’ rules and use local policies for individual and localised rules. | 4 | 2 | 2 | REFACTOR |
| Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs. | 5 | 4 | 3 | REIMPLEMENT |
| Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny. | 5 | 3 | 2 | REFACTOR |
| Configure Azure firewall IDS in Deny mode. | 5 | 2 | 2 | REFACTOR |
| Consider using IP groups to reduce IP table rules usage. | 5 | 2 | 1 | REFACTOR |
5 Monitoring & Logging
5.1 Overview
Monitoring within Azure is completed by four services; Azure Monitor, Azure Service Health, Azure Advisor and Microsoft Defender for Cloud, all of which serve a different purpose or function. Within this section we will mainly be looking at Azure Monitor as a service and its capabilities as a single unified hub for all monitoring and diagnostics data within Azure. To maximise the value out of Monitoring we need to ensure that we have a well-defined and designed Logging platform, or as the CAF calls it Monitoring Data Platform.
Figure 5‑1: Azure Monitor Service Overview
5.2 Logging
5.2.1 Definition
Logging within Microsoft Azure is comprised of three different classifications of what is referred to as “Platform Logs”. These logs are used to provide detailed information ranging from diagnostic information about specific resources to auditing information for activity logs or Azure Active Directory.
Resource Logs
Resource Logs, previously known as “Diagnostic Settings”, provide insight into operations that were performed within an Azure resource (the data plane), for example getting a secret from a Key Vault, or making a request to a database. The content of resource logs varies by the Azure service and resource type.
Activity Logs
Activity Logs, provides insight into the operations on each Azure resource within the context of a subscription, from management plane, in addition to updates on Service Health events. Activity logs can be used to determine what, who and when for any write operations (PUT, POST, DELETE) taken on the resources within your subscription.
Microsoft Entra Logs
Entra Logs, contains all the sign-in activity and audit trail of changes made in the Entra for a particular tenant.
5.2.2 Current State
CUST currently have recently implemented standards around the capturing of log data within Azure. Logging capture has been implemented into a single Sentinel enabled LAW.
It was found, however, that there are several ‘default’ workspaces that are not within the logging construct and do not follow naming conventions within the CUST environment.
5.2.3 Cloud Adoption Framework (CAF)
The CAF has three strategies that can be applied to Logging, those being:
- Centralised – All logs are stored in a central workspace and managed by a single team. With Azure Monitor providing differential access on a per-team basis. In this case, it is easy to manage, search across resource and cross-correlate logs. But one of the main limitations of this model is the administrative overheard that eventually comes as the use of the workspace grows. This model is also known as the Hub and Spoke.
- Decentralised – Each team has their own workspace created in a resource group that they can each own and manage, where log data is segregated per resource. In this case, the workspace can be kept security and access control is able to be kept consistent with resource access. But it’s extremely difficult to cross-corelate logs. Users and Enterprises who need to create a broad view of the Azure landscape will find it difficult to analyse the data in a meaningful way without large amounts of effort.
- Hybrid – It is common to see organisations attempt to deploy both above strategies in parallel which leads to a complex, expensive, and hard-to-maintain configuration that also has gaps in coverage.
5.2.4 Recommendations and Assessment
CUST should look to implement a centralised data collection / logging capability within Azure. This does not mean a singular instance of Log Analytics but keeping the number as small as possible will ensure that the environment is simple and manageable.
Whilst keeping the number of Log Analytics Workspaces to a minimum it is important to note that ingesting data across regions within Azure does come with a cost. So where required a separate Log Analytics Workspace should be created per geographical region.
Figure 5‑2: Diagnostic Settings Logging Pattern
Within the context of each Log Analytics resource there is also multiple levels of access controls that can be applied. As there has already been a large amount of upfront effort within the management of access to resources within Azure, utilising the Azure Permissions access model is recommended, this enables low operational overhead whilst enabling re-use of granular permissions declared against the Azure resources.
| Logging Recommendations |
– Consider segregating Operational logging data from Security logging data to save on cost. – Establish Logging patterns for an all services (PaaS and IaaS) and document them and deploy consistently via Azure Policy. – Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data. – Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards. – Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Consider segregating Operational logging data from Security logging data to save on cost. | 5 | 3 | 3 | REFACTOR |
| Establish Logging patterns for an all services (PaaS and IaaS) and document them and deploy consistently via Azure Policy. | 5 | 4 | 2 | REIMPLEMENT |
| Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data. | 5 | 4 | 2 | REIMPLEMENT |
| Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards. | 5 | 3 | 3 | REIMPLEMENT |
| Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure. | 5 | 2 | 2 | REIMPLEMENT |
5.3 Monitoring
5.3.1 Definition
Azure Monitor is the native Azure service to enable the maximum availability and performance potential of your applications and services by enabling a single platform for collecting, analysing, and acting on telemetry from your cloud and on-premises environments.
Azure Monitor is targeted to help you proactively identify issues within your applications and infrastructure whilst also providing insight into the performance of these services.
In August 2018, Microsoft consolidated Log Analytics, and Application Insights into the service we now know as Azure Monitor. This was done to enable a single integrated experience of monitoring Azure resources and hybrid environments.
5.3.2 Current State
CUST have implemented Azure monitoring but there is a limited capacity.
Currently, there has been no alerting configured to monitor Azure usage and events from logs captured within Log Analytics. Application Insights has been configured on 2 instances to gain low level application telemetry that is exposed into the Azure Monitor platform. Azure Dashboards are not used to assist with monitoring applications and network utilisation.
5.3.3 Cloud Adoption Framework (CAF)
Azure Monitor is the Azure native platform service that provides a single source for monitoring Azure resources. It can monitor all layers of the stack, starting with Tenant services, such as Azure Active Directory Services through to Application Insights such as SQL Queries between an Azure Web Application and an Azure SQL Database. It is because of this full stack capability that the CAF recommends utilising Azure Monitor wherever possible for monitor as opposed to 3rd party tooling. The CAF does cover Alerting recommendations for within Azure Monitor, but this is more just what you can alert on with Azure Monitor rather than best practices.
5.3.4 Recommendations and Assessment
We recommend that CUST have a documented Monitoring strategy that takes a service-orientated approach to clearly define what monitoring is to be used for the infrastructure layer, resource, scope, and the method that is to be used for monitoring. The below recommendations are what should be implemented in Azure Monitor and, any application monitoring requires definition and validation of effectiveness to ensure the end-to-end implementation meets the business requirements.
| Monitoring Recommendations |
– Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone. – Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example), alert on anomalies. – Enable VM Insights and performance counters for relevant workloads. – Configure Service and Resource health alerts and notify central platform and operational teams when issues occur. – Define and implement Monitoring for all PaaS Services. – Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications. – Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone. | 5 | 3 | 2 | REIMPLEMENT |
| Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example). | 5 | 2 | 2 | REFACTOR |
| Configure Service and Resource health alerts and notify central platform and operational teams when issues occur. | 4 | 2 | 1 | REFACTOR |
| Define and implement Monitoring for all PaaS Services. | 3 | 2 | 2 | REFACTOR |
| Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications. | 3 | 2 | 2 | REFACTOR |
| Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns. | 4 | 2 | 1 | REFACTOR |
5.4 Azure Dashboards
5.4.1 Definition
Azure Dashboards are a focused and organized view of cloud resources within the Azure Portal. Dashboards can be used as a workspace where you can quickly launch relevant tasks for your day-to-day operational activities, such as the main resource group for the application you are currently working on. Dashboards are also able to be published and shared across an Azure Tenant, to enable a focused visibility on specific environments or capabilities within Azure or even On-Premises data.
5.4.2 Current State
CUST have not implemented the use of dashboards to monitor the usage and health of any resources running on Azure. Dashboards can be created to monitor the express route utilisation and cover the standard infrastructure landscape that support the Platform to give a snapshot view of the operational state of core services.
5.4.3 Cloud Adoption Framework (CAF)
Dashboarding is not specifically mentioned within the CAF.
5.4.4 Recommendations and Assessment
Azure Dashboards are the best way to get visibility across an applications environment. As the data to make use of Azure Dashboard will already be collected there is minimal effort required to configure the use of Azure Dashboards across CUST.
| Dashboard Recommendations |
– Create the ability for Templated Dashboards / Workbooks as shareable entities. – Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications. – Extend the use of Dashboards / workbooks for Infrastructure monitoring. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Create the ability for Templated Dashboards / Workbooks as shareable entities. | 2 | 1 | 1 | REFACTOR |
| Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications. | 2 | 1 | 1 | REFACTOR |
| Extend the use of Dashboards / workbooks for Infrastructure monitoring. | 3 | 2 | 1 | REIMPLEMENT |
6 Identity
6.1 Overview
In any environment, whether on-premises, hybrid, or cloud-only, IT needs to control which administrators, users, and groups have access to resources. Identity and access management (IAM) services enable you to manage access control in the cloud.
Several options are available for managing identity in a cloud environment. These options vary in cost and complexity. A key factor in structuring your cloud-based identity services is the level of integration required with your existing on-premises identity infrastructure.
Figure 6‑1: Azure Active Directory complexity scale
6.2 Entra ID
6.2.1 Definition
Entra ID is Microsoft’s cloud-based identity and access management in Azure Cloud. It’s at the centre of authentication and authorisation for users access as well as applications and machine identities.
Entra ID uses modern authentication protocols and is fully accessible via Microsoft Graph API for programmatic accesses. Depending on the licencing model chosen, a range of features are available to enhance visibility and security controls.
A common scenario for businesses is to use Entra ID in a hybrid configuration. Apart from startups that would be Cloud native, most companies would have an existing identity solution on-premises to manage their users’ accesses to internal and external resources.
6.2.2 Current State
CUST have a single Entra ID tenant within Azure. Microsoft Entra Connect with Connect Sync has been configured to synchronise identity with on-premises domain.
| Tenant Display Name | Domain Names | Tenant ID | Subscriptions |
CUST PTY LTD
| CUST CUSTPTYLTD.onmicrosoft.com | aa391e1f-5f1c-48e3-aef0-4c1c91da1718 | Yes |
The current secure Score on Entra is 74% and the following areas of improvement should be targeted, targeting the Low Implementation cost and Low user impact first.
| Name | Score Impact | Current Score | Max Score | User Impact | Implementation Cost | Status |
| Stop clear text credentials exposure | 1.83 | 0 | 5 | Low | Low | To address |
| Modify unsecure Kerberos delegations to prevent impersonation | 1.83 | 0 | 5 | Low | Low | To address |
| Reduce lateral movement path risk to sensitive entities | 1.83 | 0 | 5 | Low | Low | To address |
| Protect and manage local admin passwords with Microsoft LAPS | 1.83 | 0 | 5 | Low | Low | To address |
| Resolve unsecure account attributes | 1.83 | 0 | 5 | Low | Low | To address |
| Stop weak cipher usage | 1.83 | 0 | 5 | Low | Low | To address |
| Change password of built-in domain Administrator account | 1.83 | 0 | 5 | Low | Low | To address |
User settings for App registrations and Administration portal are appropriate. Linkedin account connections is enabled for all users, allowing this grants access to user’s properties, confirm that this is required or not.
External collaboration settings are configured appropriately, other than the Collaboration restrictions which may need review.
CUST are leveraging named locations
6.2.3 Cloud Adoption Framework (CAF)
For organizations with existing on-premises Active Directory infrastructure, directory synchronization is often the best solution for preserving existing user and access management while providing the required IAM capabilities for managing cloud resources. This process continuously replicates directory information between Entra ID and on-premises directory services, allowing common credentials for users and a consistent identity, role, and permission system across your entire organisation.
Directory synchronization assumptions: Using a synchronized identity solution (Entra Connect) assumes the following:
- You need to maintain a common set of user accounts and groups across your cloud and on-premises IT infrastructure.
- Your on-premises identity services support replication with Entra ID.
Cloud-hosted domain services assumptions: Performing a directory migration assumes the following:
- Your workloads depend on claims-based authentication using protocols like Kerberos or NTLM.
- Your workload virtual machines need to be domain-joined for management or application of Active Directory group policy purposes.
Entra federation services assumptions: Using EFS assumes the following:
- Your workloads require a single sign-on capability cross multiple domains within an organisation.
6.2.4 Recommendations and Assessment
CUST have a good implementation of Entra ID and are using Azure managed identities as a principal type within the platforms DevOps pipelines.
| Dashboard Recommendations |
– Review Linkedin account connections being enabled for all users. – Allow invitations only to the specified domains, ensure there is a regular review of the targeted domains. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Review Linkedin account connections being enabled for all users. | 3 | 1 | 1 | REFACTOR |
| Allow invitations only to the specified domains, ensure there is a regular review of the targeted domains. | 3 | 1 | 1 | REFACTOR |
6.3 Role Based Access Control
6.3.1 Definition
Access management for cloud resources is a critical function for any organization that is using the cloud. Azure role-based access control (Azure RBAC) helps you manage who has access to Azure resources, what they can do with those resources, and what areas they have access to.
Azure RBAC is an authorization system built on Azure Resource Manager that provides fine-grained access management of Azure resources.
6.3.2 Current State
CUST have implemented RBAC roles for user access to subscription resources and Entra objects. There is a consistent implementation of RBAC roles has been observed, however, they are generic and aligned to the Azure Roles instead of an CUST Persona/Operating model within the platform and generated landing zones.
It was also noted that CUST are using Access Packages to allow Teams to request access to the relevant resources, whilst good practice the requestable roles should be Persona/Role aligned and protected with Azure ABAC to streamline the approval process and ensure platform users are not over permissioned.
6.3.3 Cloud Adoption Framework (CAF)
Only Grant the access users need:
Using Azure RBAC, you can segregate duties within your team and grant only the amount of access to users that they need to perform their jobs. Instead of giving everybody unrestricted permissions in your Azure subscription or resources, you can allow only certain actions at a particular scope.
When planning your access control strategy, it’s a best practice to grant users the least privilege to get their work done. Avoid assigning broader roles at broader scopes even if it initially seems more convenient to do so. When creating custom roles, only include the permissions users need. By limiting roles and scopes, you limit what resources are at risk if the security principal is ever compromised.
The following diagram shows a suggested pattern for using Azure RBAC.
Figure 8: Role Based Access Control Pattern
Limit the number of subscription owners:
You should have a maximum of 3 subscription owners to reduce the potential for breach by a compromised owner. This recommendation can be monitored in Microsoft Defender for Cloud.
Use Entra ID Privileged Identity Management:
To protect privileged accounts from malicious cyber-attacks, you can use Entra ID Privileged Identity Management (PIM) to lower the exposure time of privileges and increase your visibility into their use through reports and alerts. PIM helps protect privileged accounts by providing just-in-time privileged access to Entra ID and Azure resources. Access can be time bound after which privileges are revoked automatically.
6.3.4 Recommendations and Assessment
CUST have implemented a good RBAC model for the Entra tenants. Continual review of the RBAC model should be implemented to assess the organization access requirement and ensure a least privileged model is implemented to reduce the impact and risk of compromised or mishandled identity access.
| Role Based Access Recommendations |
– Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments and access packages. – Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones. – Consider establishing a cloud-based remote-access solution for Azure Administration. – Re define PIM to audit and validate that PIM escalation are tightly controlled. – Conduct PIM access reviews periodically. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments and access packages. | 5 | 4 | 4 | REIMPLEMENT |
| Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones. | 5 | 3 | 3 | REFACTOR |
| Consider establishing a cloud-based remote-access solution for Azure Administration. | 5 | 4 | 4 | REIMPLEMENT |
| Re define PIM to audit and validate that PIM escalation are tightly controlled. | 4 | 2 | 2 | REFACTOR |
6.4 Conditional Access
6.4.1 Definition
Conditional Access is the tool used by Azure Active Directory to bring signals together, to make decisions, and enforce organizational policies. Conditional Access is at the heart of the new identity driven control plane.
By using Conditional Access policies, you can apply the right access controls when needed to keep your organization secure and stay out of your user’s way when not needed.
Figure 9: Conditional Access Pattern
6.4.2 Current State
CUST currently have 28 policies enabled. Policies configured within the CUST Entra Tenant are aligned with ASD’s Blueprint for Secure cloud. At a high-level, these policies require the following standards:
Base Protection Policy
- All users and administrators required to use MFA
- Blocks legacy authentication
- Requires compliant or hybrid Entra joined devices
- Blocks access from non-trusted locations
Privileged Access Policy
- Requires MFA for all privileged role holders
- Additional device compliance requirements for privileged accounts
- Session controls including sign-in frequency
- Restricted access to administration portals
- Blocks access from non-corporate devices
Session Controls
- Sign-in frequency: 4 hours
- Persistent browser session: Disabled
- Continuous access evaluation: Enabled
- Device filter: Compliant devices only
- Application enforced restrictions for sensitive data
6.4.3 Cloud Adoption Framework (CAF)
Conditional Access is not specifically mentioned within the CAF.
The CAF references using Security defaults as a baseline. Security defaults make it easier to help protect your organization from these attacks with preconfigured security settings:
- Requiring all users to register for Microsoft Entra Multi-Factor Authentication.
- Requiring administrators to perform multi-factor authentication.
- Blocking legacy authentication protocols.
- Requiring users to perform multi-factor authentication when necessary.
- Protecting privileged activities like access to the Azure portal.
6.4.4 Recommendations and Assessment
Conditional Access Policies are configured appropriately within the CUST tenant.
6.5 Privileged Identity Management
6.5.1 Definition
Privileged Identity Management (PIM) is a service in Microsoft Entra ID that enables you to manage, control, and monitor access to important resources in your organization. These resources include resources in Entra ID, Azure, and other Microsoft Online Services such as Microsoft 365 or Microsoft Intune
Privileged Identity Management provides time-based and approval-based role activation to mitigate the risks of excessive, unnecessary, or misused access permissions on resources that you care about. Features of Privileged Identity Management:
- Provide just-in-timeprivileged access to Entra ID and Azure resources
- Assign time-boundaccess to resources using start and end dates
- Require approvalto activate privileged roles
- Enforce multi-factor authenticationto activate any role
- Use justificationto understand why users activate
- Get notificationswhen privileged roles are activated
- Conduct access reviewsto ensure users still need roles
- Download audit historyfor internal or external audit
6.5.2 Current State
PIM has been implemented across CUST Entra tenant. A Consistent implementation has been observed for privileged Entra roles.
There were 2 assignments which were listed as permanent within the Global Administrator role which seem to be break glass accounts, although they will need validation.
6.5.3 Cloud Adoption Framework (CAF)
Microsoft recommends the following Azure roles managed by Privileged Identity Management as a baseline
- Global administrator
- Security administrator
- User administrator
- Exchange administrator
- SharePoint administrator
- Intune administrator
- Security reader
- Service administrator
- Billing administrator
- Skype for Business administrator
Microsoft recommends that you manage Owner roles and User Access Administrator roles of all subscriptions/resources using Privileged Identity Management.
Microsoft recommends you manage all roles with guest users using Privileged Identity Management to reduce risk associated with compromised guest user accounts.
Microsoft recommends that you bring Entra ID role-assignable groups under management by Privileged Identity Management.
Microsoft recommends you have zero permanently active assignments for both Entra ID roles and Azure roles other than the recommended two break-glass emergency access accounts.
Microsoft recommends you set up Azure log monitoring to archive audit events in an Azure storage account for greater security and compliance.
6.5.4 Recommendations and Assessment
| Privileged Identity Management Recommendations |
– Review permanent Global Administrator assignments. – Review PIM role assignments against Microsoft PIM baseline. – Conduct regular review of PIM alerts. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Review permanent Global Administrator assignments. | 3 | 1 | 1 | REFACTOR |
| Review PIM role assignments against Microsoft PIM baseline. | 3 | 1 | 1 | REFACTOR |
| Conduct regular review of PIM alerts. | 4 | 1 | 1 | REFACTOR |
7 Security
7.1 Defender for Cloud
7.1.1 Definition
Microsoft Defender for Cloud is a unified infrastructure security management system that strengthens the cloud security posture and provides advanced threat protection across your hybrid workloads in the cloud – whether they are in Azure or not – as well as on-premises. Microsoft Defender for Cloud comes with features built-in such as:
- Workflow Automation: backed by Azure Logic Apps which allow you to automate responses to certain alerts e.g., send email notifications and raise a ticket in Service Now.
- Azure Policies and Compliance: Visibility and real-time reporting on your security posture according to predefined or custom Azure Policies.
- Advanced Cloud Defence:
- Just-in-time virtual machine access
- Opens public IP to allow for RDP/SSH connection (Updates NSG on VM)
- Disable recommendations around JIT access as it contradicts CIS and Public endpoints
- Adaptive application controls
- Automated intelligence framework that helps reduce attack service to VMs
- Create alerts if unknown or trusted applications run on VM that has not been whitelisted.
- Threat Protection: provides comprehensive defence for your environment such as Azure compute resources (VMs, App Services, Containers), data resources (Azure SQL, CosmosDB) and service layers (VNet, Key Vaults).
- Just-in-time virtual machine access
7.1.2 Current State
Microsoft Defender for Cloud is broken up into 4 main areas of management those being:
- Policy & Compliance
- Resource Security Hygiene
- Threat Protection
- Advanced Cloud Defence
As it currently stands CUST are consuming the Azure Defender CSPM plan in partial workload protection configuration in all the onboarded subscriptions. There is a consistent configuration across all the subscriptions and plans with a good coverage of resources.
The posture within Policy and Compliance is based on the Microsoft “Secure Score” which is calculated based on the ratio between your healthy resources and your total resources. A “Healthy” resource is determined by any recommendations that are generated from the Microsoft Defender for Cloud. CUST’s current Secure Score across its subscriptions are at the time of publishing this assessment:
Threat protection provides insight into perceived threats within your azure environments specifically alerting on issues that are against Azure Compute, Data or Service layer resources. These alerts are usually paired with recommended remediation steps but in some instances can also be used to trigger downstream actions.
There is 1 active alert were observed within Threat Protection
There are several security recommendations noted within Defender, some labelled as critical and high which require intervention.
7.1.3 Cloud Adoption Framework (CAF)
The Cloud Adoption Framework largely focuses on the use of Defender for cloud to pre-emptively detect vulnerabilities across the Azure environment. The recommendations from the CAF are directly aligned to the recommendations that can be seen within the Defender for Cloud as it looks to assess the deployed resources within your Azure environments. This includes things such as Just-in-time (JIT) access implementation for Azure Virtual Machines.
7.1.4 Recommendations and Assessment
CUST have done a comprehensive job collecting and feeding data Defender for Cloud by enabling data collection for Azure resources to Log Analytics. To take advantage of this effort, a review & action of the recommendations active in Defender to further secure azure resources where appropriate.
| Defender for Cloud Recommendations |
– Investigate dismissing false positive recommendations within Defender. – Establish a weekly review of Defender recommendations. – Review the Regulatory Compliance reporting within Defender and consider the recommended controls, some of which require Azure policy enforcement or configuration. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Investigate dismissing false positive recommendations within Defender. | 4 | 1 | 1 | REFACTOR |
| Establish a weekly review of Defender recommendations. | 4 | 2 | 1 | REFACTOR |
| Review the Regulatory Compliance reporting within Defender and consider the recommended controls, some of which require Azure policy enforcement or configuration. | 4 | 2 | 2 | REFACTOR |
7.2 Azure Sentinel
7.2.1 Definition
Microsoft Sentinel is a scalable, cloud-native, Security Information Event Management – SIEM. Azure Sentinel delivers intelligent security analytics and threat intelligence across the enterprise, providing a single solution for alert detection, threat visibility, proactive hunting, and threat response.
Sentinel provides a correlation point for alerts and information from all Microsoft services.
7.2.2 Current State
CUST currently leverage Azure Sentinel out of the central Log Analytics workspace.
There are several watchlists in place but could be extended to detect sensitive accounts & break-glass account compromise.
7.2.3 Cloud Adoption Framework (CAF)
Whilst there are no CAF specific recommendations, Microsoft’s guidance is that CUST should define and document clear security objectives and asset identification, followed by establishing governance through RBAC and retention policies. Security operations benefit from automated responses and cross-workspace monitoring capabilities. The management approach requires regular cost monitoring, alert tuning, and archival strategies. Operational effectiveness depends on logical resource organisation and health monitoring of data sources. These practices work together to create a comprehensive security information and event management (SIEM) solution that supports both security and business objectives.
7.2.4 Recommendations and Assessment
Azure Sentinel has an extremely easy integration model to Azure services, Enabling some Sentinel’s default capabilities would provide additional security insight to potential security issues within CUST’s Azure environment.
| Sentinel Recommendations |
– Enable UEBA to detect and Identify Insider threats in CUST’s environment and their potential impact. – Review and extend the use of Microsoft Sentinel watchlists to investigate threats, import data, reduce alert fatigue & enrich event data derived from multiple data sources. – Refer to the Microsoft community Playbooks on Github for other security-related playbooks as per business requirements. – Re-validate enabled \ disabled playbooks for response to incidents. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Enable UEBA to detect and Identify Insider threats in CUST’s environment and their potential impact. | 4 | 1 | 2 | REFACTOR |
| Review and extend the use of Microsoft Sentinel watchlists to investigate threats, import data, reduce alert fatigue & enrich event data derived from multiple data sources. | 4 | 2 | 1 | REFACTOR |
| Refer to the Microsoft community Playbooks on Github for other security-related playbooks as per business requirements. | 3 | 1 | 2 | REFACTOR |
| Re-validate enabled \ disabled playbooks for response to incidents. | 4 | 2 | 2 | REFACTOR |
8 DevOps
8.1 Definition
DevOps is the union of people, processes, and technology to continually provide value to customers. DevOps enables formerly siloed roles: development, IT operations, quality engineering and security to coordinate and collaborate to produce better, more reliable products.
By adopting a DevOps culture along with DevOps practices and tools, teams gain the ability to better respond to customer needs, increase confidence in the applications they build and achieve business goals faster.
8.2 People Process and Culture
8.2.1 Definition
The Cloud Adoption Framework identifies Devops only in terms of building a Cloud Adoption Plan.
You can’t buy DevOps, as DevOps is not a software, tool, process, company, or person, it’s a methodology used especially by IT professionals.
DevOps is the correlation of people, process, and products to enable continuous delivery of value to end users. The outcomes are tightly connected to allow for frequent releases and at the same time to keep the same level of quality.
People
Initially, stakeholders would need to be identified, ensuring that everybody delivering value to the business are working tightly together on the common goal of adding value to the customer. The latest study by Gartner, and from this study we can deduce that the biggest concern is ‘People’, with process and products being deemed less critical. From this, we can assume that having highly motivated people with good collaboration is necessary.
Process
Next – DevOps is about improving process because even if you have highly motivated people that are working well together, you may still have several business processes which may get in the way, and this could really block innovation. For instance, having to seek approval from long chain advisory boards before implementing changes or being restricted to doing things in a certain way can impede innovation. The process of designing, building, and testing software should be well presented to each individual team member, making them aware of all parts of the development process. The implementation of DevOps can be hard work, as it completely changes the company’s structure.
The core of this is enabling efficient flow of collaboration and making sure that business structure and processes do not get in the way but instead have processes and practices that help improve the value and delivery to your customers.
Products
Products, tools, and services that can help enable different DevOps practices and different teams which can be used to make things easier. From a very high level, these tools include Microsoft Azure, which offers a lot of different products and services.
Figure 11: Sample Mapping of skills to IT roles in a cloud-hosted environment.
The following table represents the suggested RACI per Team but also how each team will interact with one another.
| Team | Solution delivery | Business alignment | Change management | Solution operations | Governance | Platform operations | Platform automation | ||
| Strategy team | C | A | A | C | C | I | I | ||
| Adoption team | A | C | R | C | I | I | I | ||
| Ops team | C | C | R | A | C | A | C | ||
| Gov team | C | I | I | C | A | R | I | ||
| Aligned Capability | Cloud adoption | Cloud strategy | Cloud strategy | Cloud operations | CCoE & Governance | CCoE and Cloud platform | CCoE and Cloud automation | ||
8.2.2 Current State
CUST have a dedicated DevOps team that works closely alongside key Infrastructure and Security teams to deliver platforms and solution.
CUST currently leverage various teams across the business that are autonomous from each other but still collaborate on joint project where required. Responsibility for Azure infrastructure mostly falls to the cloud Infrastructure and Security chapters.
Within the DevOps team, there are highly skilled engineers but it is not clear what functions they perform within the team and platform. There is clear deliniation between Platform and Landing zones but there is no clear definition that same structure in relation to personas and functions so it is not clear to what degree onboarding teams will be enabled and autonomous.
There does not appear to be a defined Enterprise approach to DevOps and architecture so the strategy is loosley defined which may result in pcokets of high and low skill areas within the organistaion. Sprints within the DevOps team occur monthly or every 4 weeks and there are no retrospectives captured to understand if adequate progress has been made against the backlog of work.
8.2.3 Cloud Adoption Framework (CAF)
The Azure CAF generally talks to the uptake of services in this space and how organisations can flex and change to allow for this.
- Digital estate rationalisation:What are the top 10 priority workloads in the adoption plan? How many additional workloads are likely to be in the plan? How many assets are being considered as candidates for cloud adoption? Are the initial efforts focused more on migration or innovation activities?
- Organisation alignment:Who will do the technical work in the adoption plan? Who is accountable for adherence to governance and compliance requirements?
- Skills readiness:How many people are allocated to perform the required tasks? How well are their skills aligned to cloud adoption efforts? Are partners aligned to support the technical implementation?
8.2.4 Recommendations
| People, Process, Technology Recommendations |
– Establish and document enterprise DevOps standards. – Define the RACI matrix for current teams to identify and address any existing silos. – Create a CoE forum/framework to facilitate team collaboration without management or process barriers, promoting shared ownership. – Use the forum to iteratively enhance services with transparent feedback for seamless improvements and adoption. – Centralise ownership of targeted IaC components/modules and implement a feedback loop for overall improvements (e.g., distributed commit, central creation). – Formalise the operational support approach and align capabilities to ensure operational resilience. – Reduce sprint time to 2 week sprints and establish a process to perform restrospectives on delivered sprint to capture learnings and imrpove. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Establish and document enterprise DevOps standards. | 5 | 3 | 5 | REIMPLEMENT |
| Define the RACI matrix for current teams to identify and address any existing silos. | 5 | 3 | 2 | REIMPLEMENT |
| Create a CoE forum/framework to facilitate team collaboration without management or process barriers, promoting shared ownership. | 5 | 3 | 2 | REIMPLEMENT |
| Use the forum to iteratively enhance services with transparent feedback for seamless improvements and adoption. | 5 | 2 | 2 | REIMPLEMENT |
| Centralise ownership of targeted IaC components/modules and implement a feedback loop for overall improvements (e.g., distributed commit, central creation). | 5 | 3 | 3 | REIMPLEMENT |
| Formalise the operational support approach and align capabilities to ensure operational resilience. | 5 | 4 | 3 | REIMPLEMENT |
| Reduce sprint time to 2 week sprints and establish a process to perform restrospectives on delivered sprint to capture learnings and imrpove. | 5 | 2 | 1 | REIMPLEMENT |
8.3 Architecture
8.3.1 Definition
The following section includes a list of specific Architecture artefacts to be created/modified to include automation / DevOps capabilities and their associate interactions across the enterprise.
These artefacts will be critical in driving rationalisation, & integration objectives.
Strategy & Vision
The intent of the DevOps and Automation strategy is to define the guiding principles and characteristics of the associated roadmaps, blueprints, and reference architectures. The architecture team should utilise the governance structures defined in later sections, along with direct business and wider chapters within CUST.
Enterprise Architecture
To support the reuse, consolidation and efficiency potential of DevOps, there must be a targeted effort in defining the business and technical interfaces.
In order achieve the goal of repeatable end to end build pipelines of business applications, the storage, sharing, and control of application code, data and configuration information should be defined up front and not as a by-product of early DevOps / Cloud adoption efforts.
Application Architecture
The application architecture of business applications should be defined before or in the very early stages of DevOps onboarding or cloud migration / build. Clear guidance of the intended architecture will influence hosting, build, validation, and branching strategies of not just the application in scope but it’s child, parent, and other applications in interacts / shares with. Clearly articulating the application architecture will allow Infrastructure and Operations teams to better shape the establishment of the Services and Systems for consumption by applications teams. Over time, as the hosting reference architecture and service catalogues of the new contemporary platforms matures it will in part begin to influence the application architecture in bidirectional way.
Reference Architecture and Blueprints
A series of reference architecture & blueprint artefacts should be completed for high priority / cloud-based hosting patterns. This will allow the development of deployment pipelines with the necessary building blocks end to end build pipelines.
8.3.2 Current State
From a Platform perspective, the Architectural and design documentation does not describe the assessed environment accurately. Within the design documenation that was assessed, there is limited detail and an unsutable level of context which in turn makes it difficult the understand for Architecture, and difficult to build too for an engineering team. These discrepincies generally lead to a substandard output and an increase in technical debt and refactoring post build. This mis-alignment makes it challenging to communicate the current state to new Team members or to share with outside teams that will be consumers of this platform, placing greater pressure and reliance on the Cloud DevOps team.
There is currently no defined process for defining architectural planning workflows or documentation of CUST’s Azure tenant or any solutions deployed to it. During the discovery workshop, it was described as mostly an ad-hoc depending on the solution and the chapter that oversees it.
8.3.3 Cloud Adoption Framework (CAF)
The Cloud Adoption Framework enterprise-scale landing zone architecture represents the strategic design path and target technical state for an organization’s Azure environment. It will continue to evolve alongside the Azure platform and is defined by the various design decisions an organization must make to map an Azure journey.
Not all enterprises adopt Azure the same way, so the Cloud Adoption Framework enterprise-scale landing zone architecture varies between customers. The technical considerations and design recommendations in this guide might yield different trade-offs based on your organization’s scenario. Some variation is expected, but if you follow the core recommendations, the resulting target architecture will set your organization on a path to sustainable scale.Cloud foundations workshop, it was described as mostly an ad-hoc depending on the solution and the chapter that oversees it.
8.3.4 Architecture Reccomendations
| Architecture Recommendations |
– Document Architecture Strategy and Vision as set of guiding principles / technical charter. – Identify and document Information Architecture to the treatment of data sensitivity. – Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables). – Uplift current Platform Architectural documentation to accurately reflect the implemented platform. – Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework. – Create solutions module repository of common reference Terraform modules for deploying Azure resources. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Document Architecture Strategy and Vision as set of guiding principles / technical charter. | 5 | 3 | 3 | REIMPLEMENT |
| Identify and document Information Architecture to the treatment of data sensitivity. | 5 | 3 | 2 | REIMPLEMENT |
| Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables). | 5 | 2 | 2 | REIMPLEMENT |
| Uplift current Platform Architectural documentation to accurately reflect the implemented platform. | 5 | 4 | 2 | REIMPLEMENT |
| Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework. | 5 | 3 | 3 | REIMPLEMENT |
| Create solutions module repository of common reference Terraform modules for deploying Azure resources. | 5 | 3 | 2 | REIMPLEMENT |
8.4 Azure DevOps
8.4.1 Definition
Azure DevOps is Microsoft’s SaaS offering for a holistic DevOps toolchain. It provides developer services to support teams to plan work, collaborate on code development and build/deploy applications and infrastructure. Azure DevOps is comprised of 5 core service functions:
| Service Name | Description |
| Azure Repos | Provides Git repositories or Team Foundation Version Control (TFVC) for source control of your code |
| Azure Pipelines | Provides build and release services to support continuous integration and delivery of your apps |
| Azure Boards | Delivers a suite of Agile tools to support planning and tracking work, code defects, and issues using Kanban and Scrum methods |
| Azure Test Plans | Provides several tools to test your apps, including manual/exploratory testing and continuous testing |
| Azure Artifacts | Allows teams to share Maven, npm, and NuGet packages from public and private sources and integrate package sharing into your CI/CD pipelines |
8.4.2 Current State
CUST have implemented Azure DevOps with the One Organisation, Many Projects, Many Teams principle, which is the implementation model recommended, this both reduces significant operational overhead from a platform teams’ perspective whilst balancing the most enablement for application/development teams.
The DevOps team will be responsible for Landing Zone build which means each onboarding appplication/team will have their own Project, Service connection and Managed identity to deploy infrastructure, which creates a clear delineation between platform and application teams.
Platform repo’s make use of feature and fix branches for both complex changes and there is a branch policy on the ‘Master’ branch which requires 1 approval to merge, however, the policy does allow requestors to approve their own changes. The policy also has comment resolution checks enabled, however, comment resolution is optional meaning that pull requests could be commited with unresolved comments.
In the context of automation, the DevOps pipelines and Terraform code leverage a storage account as the backend for all Platform state.
Whilst this is good practice, there are a few configurations that require changes to have this particular storage account in a best-practice state:
- Configure ABAC rules on a container level to only allow access to .tfstate files
- Disable SAS keys and access tokens, enforcing Entra ID and RBAC only
- Enable Blob versioning
- Extend soft-delete to 30 days
- Enable last accesss tracking time
8.4.3 Cloud Adoption Framework (CAF)
The CAF has a dedicated section for the use of Platform automation and DevOps, covering how most traditional IT operating models are not compatible with the cloud. Requiring an operational and organisational transformation to delivery against what are most likely significant enterprise migration targets. It is there for recommended to use a DevOps based approach for both the delivery of applications and for that of a central IT capability.
Landing Zones are a concept introduced via the Cloud Adoption Framework, defined as, the output of a multi-subscription Azure environment that accounts for scale, security, governance, networking, and identity boundaries. Essentially creating a repeatable deployment space for applications and infrastructure meeting pre-defined enterprise standards.
Outlined within the CAF are definitions of responsibilities for cross-functional DevOps teams such as Platform, Security, Networks and Application. In addition to this there is importantly a definition of central and federated responsibilities for the application and central platform teams. It is important to ensure that the application teams are empowered to control aspects of the application as part of a shared responsibility model.
8.4.4 Azure DevOps Recommendations
Whilst Azure DevOps is a product, it is important to note that the creation of a cloud-native operating model is required to take full advantage of any investment made towards an SRE or DevOps capability. Without such an operating model, teams are often hit obstacles related to lack of empowerment from senior leadership as there is no define cadence or rhythm.
Microsoft provide a design guide for the 3 main implementations of Azure DevOps:
- One Organisation, One Project, Many Teams
- One Organisation, Many Projects, Many Teams
- Many Organisations, Many Projects, Many Teams
CUST have implemented the One Organisation, Many Projects, Many Teams, which is the implementation model I would recommend, this both reduces significant operational overhead from a platform teams’ perspective whilst balancing the most enablement for application / development teams.
Azure DevOps Pipelines can apply granular levels of RBAC controls based on a directory structure. To maximise the benefit of hierarchical inheritance this should be taken advantage of within scoping application deployments. This will further reduce the operational overhead of the central platform team managing Azure DevOps but also open opportunity for empowerment to scoped application or project teams.
| Azure DevOps Recommendations |
– Define architecture of Landing Zones in Alignment with Reference Architectures. – Reconfigure the Master branch policy to require additional approvals and enforce comment resolution. – Reconfigure state storage account to be in line with the best-practice configuration. |
| Recommendation | Benefit | Effort | Complexity | Reimplement/ Refactor |
| Define architecture of Landing Zones in Alignment with Reference Architectures. | 5 | 2 | 2 | REFACTOR |
| Reconfigure the Master branch policy to require additional approvals and enforce comment resolution. | 5 | 1 | 1 | REFACTOR |
| Reconfigure state storage account to be in line with best-practice configuration | 5 | 1 | 1 | REFACTOR |
9 APPENDIX
9.1 Service Enablement
The Service Enablement Framework for Landing Zones on Microsoft Azure helps organisations in the regulated sectors meet compliance and security requirements while accelerating digital transformation. This framework provides a structured approach to defining, mapping, and enforcing necessary controls, balancing business needs with regulatory demands. It includes a prescriptive architecture and implementation plan, leveraging Azure services to streamline the process and ensure compliance with standards like PCI DSS, NIST 800-53, and SOC 1, 2, 3.
The framework outlines an operating model with clear separation of duties, involving Platform DevOps for operationalising the Azure platform and DevOps/AppOps for managing workloads within landing zones. Key functions include system operations, automation, and control mapping, all aimed at enabling broad adoption of Azure services while maintaining compliance and security. The steps involved are:
- Define and map necessary controls for compliance and security.
- Balance business needs with regulatory demands.
- Implement a clear separation of duties between Platform DevOps and DevOps/AppOps teams.
- Utilize Azure services to streamline the enablement process.
- Ensure compliance with standards e.g. PCI DSS, NIST 800-53, and SOC 1, 2, 3.
- Follow a prescriptive architecture and implementation plan.
- Map controls to Azure services and enforce them.
- Differentiate between controls managed by Microsoft and those managed by platrform.
- Map regulatory requirements to specific Azure services and controls.
- Design the architecture to meet compliance and security needs.
- Implement the designed architecture and controls.
- Collect and maintain evidence of compliance.
- Provide examples to illustrate the implementation process.
- Outline the next steps for ongoing compliance and security management.
These steps ensure a comprehensive approach to service enablement, helping financial services organizations leverage Azure while maintaining stringent compliance and security standards.
To achieve this, the below images gives a high-level approach to initiating a request for an Azure service, Developing controls and managing ongoing requests using a defined process:
Service Enablement Framework – Microsoft
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.