Skip to content

Elementor Page #2146

PDF Loading…

Contents

1    Introduction   1

1.1   Overview.. 1

1.2   Document Purpose. 1

1.3   Objectives and Requirements. 1

1.4   Document Scope. 1

1.5   Target Audience. 3

1.6   Document Conventions. 3

2    Recommendations  4

2.1   Summary. 4

2.2   Prioritised Focus Areas. 4

3    Azure Hierarchy  8

3.1   Overview.. 8

3.2   Management Groups. 8

3.3   Subscriptions. 12

3.4   Resource Groups. 16

3.5   Policy. 17

3.6   Naming and Tagging Standards. 20

3.7   Cost Management 22

4    Network  26

4.1   Overview.. 26

4.2   Virtual WAN.. 26

4.3   Virtual Networks. 29

4.4   Private Endpoints. 32

4.5   Azure Firewall 34

5    Monitoring & Logging  36

5.1   Overview.. 36

5.2   Logging. 37

5.3   Monitoring. 41

5.4   Azure Dashboards. 43

6    Identity  45

6.1   Overview.. 45

6.2   Entra ID.. 45

6.3   Role Based Access Control 50

6.4   Conditional Access. 52

6.5   Privileged Identity Management 54

7    Security  57

7.1   Defender for Cloud. 57

7.2   Azure Sentinel 59

8    DevOps  63

8.1   Definition. 63

8.2   People Process and Culture. 63

8.3   Architecture. 67

8.4   Azure DevOps. 70

9    APPENDIX   1

9.1   Service Enablement 1

1            Introduction

1.1          Overview

RLC are conducting a top-down maturity assessment of the CUST Azure and DevOps implementation to ensure alignment with both the Microsoft Cloud Adoption Framework (CAF) and industry best practice and to assess the Landing Zones readiness to support future workload migrations.

RLC have compiled this document after working closely with key business, technical and security stakeholders to ensure current state, current challenges, and future aspirations of CUST implementation and maturity journey.

1.2          Document Purpose

The purpose of this document is to highlight the deltas between the current state of CUST’s Azure implementation and Microsoft’s Cloud Adoption Framework (CAF), whilst also ensuring industry standards and best practice.

This document will provide a current state view of each component within the scope of this document and when necessary, a proposed future state with a recommendation. This document will remain at a relatively high-level of detail to maintain the correct context of a maturity baseline. This is not a documented full CAF assessment and review but covering several key components.

1.3          Objectives and Requirements

The objective of the Azure Governance, Foundations, Cost Management & DevOps review and recommendations is determining current state maturity, along with series of recommendations to be considered and executed as a program of work to further implement best practices and strategies across the department.

Outcomes of the assessment aim to improve platform reliability, operational compliance and improve adoption for consuming teams.

1.4          Document Scope

The scope of this document is limited to the agreed scope that is outlined under the RLC-CUST-5368 statement of work.

1.4.1      In Scope

The following list details the items that are within the scope of this document.

It is important to note that whilst the CAF talks about people, process, and technology this assessment is target on the technology layer of the framework.

  • Azure Foundational Design
  • Azure Hierarchy
  • Management Group Structure
  • Subscription Structure and Decision Framework
  • Resource Group Structure
  • Policy
  • Naming and Tagging Standards
  • Cost Management
  • Networking
  • Virtual WAN
  • Virtual Network
  • Private Endpoints
  • Firewalling
  • Security
  • Governance
  • Policy
  • Compliance
  • Identity
  • RBAC
  • Conditional Access
  • Privileged Identity Management
  • Monitoring
  • Logging
  • DevOps

At the time of publication of this assessment access was limited to the following Azure tenants & subscriptions:

Tenants:

  • CUST PTY LTD (CUSTptyltd.onmicrosoft.com)

1.4.2      Out of Scope

The following list details items that are specifically out of scope of this document:

  • Creation of any resources
  • Code quality and effectiveness assessment
  • Operational Management
  • Assessment of the On-Premises Environment
  • Implementation of any recommendations
  • Tenants & Subscriptions not listed in scope.

1.5          Target Audience

  • RLC Solution Consultant – the requirements will be reviewed by the solution consultant to ensure that it is consistent with Microsoft’s Cloud Adoption Framework (CAF) and industry best practice.
  • CUST key stakeholders – to help shape and formulate the assessment against key requirements gathered during workshops.
  • CUST Compliance, Security and Cloud engineers and other technical personnel.

1.6          Document Conventions

Recommendation
Recommendations will be captured in “Recommendation” boxes throughout the document. Each of these represents a recommendation from the team for future use and adoption of the solution but not delivered as part of the scope of this document.

To consider the priority of the recommendations included in this section, the following simple structure rating of 1 to 5 (1 being Lowest and 5 being Highest) this provides an overview of each recommendation.

  • Benefit:
    • 1 = Low Business / Operational Benefit
    • 5 = High Perceived Business / Operational Benefit (or Timesaving)
  • Effort:
    • 1 = 5 Days
    • 5 = Effort will take 3 months +
  • Complexity:
    • 1 = Simple change no dependencies
    • 5 = Highly complex change with multiple dependencies.
RecommendationBenefitEffortComplexity
Network: Implement Azure VWAN534

2            Recommendations

Recommendations in this section are consolidated from the detailed review sections that follow.

Specific identification has been made for uplifts that need to occur to improve the Azure platform at CUST and the Effort verses Benefit for these changes.

As an example, a reimplementation to the service is a redeployment, a refactoring to a service is a change/reconfiguration to an existing service that can be re-integrated with the existing implementation without wholesale change.

2.1          Summary

Our assessment of the Azure environment confirms it provides a suitable foundation for CUST’s current and planned workload requirements. Through our evaluation, we examined several core technical domains to assess the platform’s alignment with Microsoft landing zone best practices and its capability to support future workload migrations.

The environment demonstrates fundamental controls and architectural elements that enable secure and scalable operations. Our review covered Identity and Access Management, Resource Organisation and Standards, Network Topology and Connectivity, Management, Governance, Platform Automation and DevOps, and Security.

Throughout this report, we detail our findings across these domains and identify specific areas where the environment can be enhanced to increase operational maturity and better align with Microsoft best practices. These opportunities for improvement range from tactical adjustments to strategic initiatives that will strengthen the platform’s capabilities.

The subsequent sections provide detailed analysis of each domain, including current state assessment, alignment with best practices, and specific recommendations for improvement. These recommendations are presented with consideration for both immediate operational needs and long-term strategic objectives.

Our assessment indicates that while the platform is ready for use, implementing the recommended improvements will further enhance its security, scalability, and operational efficiency. There is, however, a key callout noting that there are no Architectural artefacts that accurately describe the built platform and landing zone patterns.

2.2          Prioritised Focus Areas

The following pages of this document present detailed observations and recommendations about CUST’s Azure cloud environment.  From all those recommendations, the top focus areas are listed in the table below.

PrioritySectionFocus AreaComment
1Azure HierarchyResource Organisation

·         Simplify the Azure scaffold structure by grouping management groups more efficiently e.g. Business units or Functions.

·         Configure a ‘default’ management group for new subscriptions so that they are not created under the Root management group.

·         Align the enterprise agreement department and accounts to reflect the relevant business structure.

·         Enable cost management by leveraging budgets and budget alerts to track costs.

2ManagementLogging

·         Centralise logging (and enable consistently).

·         Consider segregating Operational logging data from Security logging data to save on cost.

·         Establish Logging patterns for all services (PaaS and IaaS), document and deploy consistently via Azure Policy.

·         Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data.

·         Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards.

·         Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure.

3OperationsMonitoring

·         Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone.

·         Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example).

·         Configure Service and Resource health alerts and notify central platform and operational teams when issues occur.

·         Define and implement Monitoring and alerting on all PaaS Services.

·         Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications.

·         Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns.

2NetworkTopology,  Connectivity and Security

·         Review the current Network architecture and consider uplifting to Azure Virtual WAN to ensure a future proof scalable model and simplify routing.

·         Ensure NSG’s are applied to all subnet and enforced with Azure Policy to support granular rulesets.

·         Consider the use of Security Admin Rules to enforce mandatory network ACL’s.

·         Secure Core network from standard access to improve Security.

·         Enable connectivity monitors with Alerts on key network resources.

·         Enforce the use of Private Endpoints on supported services using Azure policy.

·         Periodically review traffic analytics to ensure   network flows and traffic are as expected and manage the exceptions.

·         Leverage a Hierarchical approach to applying Azure firewall policies (Use A Global policy for Core rules).

·         Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny.

·         Configure Azure firewall IDS in Deny mode.

·         Consider using IP groups  to reduce IP table rules usage.

3Foundations

Architecture

and

Standards

·         Align design and as-built documentation to the current platform architecture and state.

·         Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework.

·         Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables).

·         Create solutions module repository of common reference Terraform modules for deploying Azure resources.

4Governance

Governance

and Cost

Management

·         Adopt a Policy driven approach to enforcing guardrails for Azure resources.

·         Leverage Azure policy to govern which Azure services are allowed to be deployed by Application teams.

·         Create Azure resource  locks to prevent accidental deletion of critical resources, these can be managed as a stage in the DevOps deployment pipelines.

·         Create a Show Back / Charge Back model for cost

management (enable a FinOps Approach)

·         Create a process to manage (and delete) Orphaned Disks and NICs

·         Ensure cost definitions are built into the inception of a services creation.

·         Action Azure Advisor recommendations where it makes sense.

·         Action Azure Advisor insights.

·         Consolidate Logging then apply Capacity on these spaces

4Identity and Access ManagementRole Based Access and Authentication

·         Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments.

·         Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones.

·         Consider establishing a cloud-based remote-access solution for Azure Administration.

·         Re define PIM to audit and validate that PIM escalation is tightly controlled.

·         Conduct PIM access reviews periodically.

6Automation and DevOpsArchitecture and Azure DevOps

·         Define the DevOps standards and principles your teams must follow.

·         Align the DevOps strategy with the business objectives.

·         Develop and define key metrics for the DevOps team to improve deployment related aspects of the platform.

·         Evaluate the effectiveness of the current DevOps practices and toolchain in relation to the Enterprise Landing Zone.

·         Determine and document how the Landing Zone team should collaborate with other teams within the organisation to effectively capture requirements and iterate on design and the platform.

·         Develop a centralised reusable set of version-controlled blueprint or modules for high-use archetypes.

3            Azure Hierarchy

3.1          Overview

Figure 3‑1: Azure Hierarchy

Azure provides four levels of management scope: those being Management Groups, Subscriptions, Resource Groups and Resources. Ensuring there is a defined structure to this management scopes is critical to managing and securing your cloud environments whilst also ensuring maximum visibility into the use and cost of your resources in Azure.

3.2          Management Groups

3.2.1      Definition

Management Groups are containers that help you manage access, policy, and compliance across multiple subscriptions. All subscriptions in a management group automatically inherit the conditions applied to the management group.

3.2.2      Current State

CUST are currently using management Groups with the structure outlined in the below image.

There is clear delineation between the Platform and Landing Zone management groups. New Landing Zones will be built under the ‘mg-application’ management group, with the depicted structure of a management group and subscription per application. The ‘mg-platform’ management group contains sub management groups for core services which are shared by all onboarding Landing Zones.

There are distinct management groups for security resources, enabling clear segregation and sandbox environments, which will be leveraged for PoC (Proof of Concept) activities.

Figure 3‑2 – Current Management Group Structure

3.2.3      Cloud Adoption Framework (CAF)

Microsoft’s Cloud Adoption Framework provides several examples of grouping subscriptions within management groups as well as proposing the use of a “Mixed” strategy for complex organisations which combine one or two of the primary strategy examples. Below is a table of these several examples:

Strategy NameDescription
Workload SeparationSubscriptions are categorised by their environment type (Production or Non-Production/Pre-Production) and placed into management groups that reflect this.
Application CategorySubscriptions are again categorised by environment type although there are usually more environment types in this approach. (Production, Development, QA etc.). There are also subscriptions created for specific workloads, such as Mission Critical or Protected Data.
FunctionalSubscriptions are categorised by functional streams such as finance, sales or IT and are placed into management groups to reflect this.
Business UnitSubscriptions are logically groups according to the business units they belong to. Business units can be defined by profit and loss category, division, profit centre or business structures.
GeographicSubscriptions are grouped within management groups based on their geographic distribution, with a management group representing each geographic location.
MixedMixed strategies comprise any two (or more) of the above 5 mentioned strategies together to define a more complex but granular hierarchy.

3.2.4      Recommendation and Assessment

CUST are currently using Management Groups and have a good implementation mostly aligned with best-practice. Although CUST doesn’t have a formalised Cloud Operating Model, the current setup is most closely aligned to the Centralised operating model of the Azure tenant by the Infrastructure and Security chapter. The current management structure supports this, but it could also support scaling to an Enterprise operating model in line with CUST’s cloud adoption growth ambitions. The Cloud Operating model should be reviewed as part CUST’s Cloud strategy. More on that can be found here.

The Applications Management Group should be renamed to ‘mg-landing-zones’ to reflect the purpose and remove the application specific Management Groups. If required, introduce another Management Group layer under ‘Landing Zones’ with an appropriate grouping e.g. Business Unit.

In the future, the platform team may consider an Archive management group to store offboarded landing zones where there is a requirement to maintain the state or data for a period after decommissioning.

Figure 3‑3: Recommended Management Group Structure

Figure 3-3 illustrates the recommended management group structure.

Management Group Recommendation

–       Decommission the management group per application structure and introduce a landing zone centric approach.

–       If an additional layer is required, introduce it as per the ‘Optional’ area depicted above.

–       Formally document the cloud operating model.

–       In the future, introduce the Archive/Decommissioned management group so that landing zones that offboarded have a holding zone in case of any retention purposes.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Decommission the management group per application structure and introduce a landing zone centric approach.421REFACTOR
If an additional layer is required, introduce it as per the ‘Optional’ area depicted above.421REFACTOR
Formally document the cloud operating model.531REFACTOR
In the future, introduce the Archive/Decommissioned management group so that landing zones that offboarded have a holding zone in case of any retention purposes.311REFACTOR

3.3          Subscriptions

3.3.1      Definition

Subscriptions logically associates user accounts and the resources that were created by those user accounts. Each subscription is subject to limits and quotas on the number of resources that can be created or used. Organisations can use subscriptions to manage costs and the resources that are created by users, teams, or projects.

3.3.2      Current State

CUST have created implemented the following subscriptions based on application environments largely based on a clear structured naming standard.

Subscription tagging has been implemented and follows a consistent and concise structure, with inheritance which propagates tags onto resource groups.

Subscription NameResource Group LocationsSecure Score
CUSTNP-Azure-TestAustralia East
CUSTNP-CAMMSNone
CUSTNP-Connectivity

Southeast Asia

Australia East

CUSTNP-DeltekNone
CUSTNP-IntegrationAustralia East
CUSTNP-JiraNone
CUSTNP-Management

Southeast Asia

Australia East

36%
CUSTNP-PredictNone
CUSTNP-PrimaveraNone
CUSTNP-Sandpit

East US

Southeast Asia

Australia East

Australia Southeast

62%
CUSTNP-SecurityAustralia East20%
CUSTPD-CAMMSNone
CUSTPD-Connectivity

Australia East

Australia Southeast

84%
CUSTPD-DeltekNone
CUSTPD-FabricAustralia East20%
CUSTPD-HybridARCAustralia East20%
CUSTPD-IdentityAustralia East
CUSTPD-Integration

Australia East

Australia Southeast

CUSTPD-InternetAustralia East39%
CUSTPD-JiraNone
CUSTPD-Management

Southeast Asia

Australia East

Australia Southeast

53%
CUSTPD-NetworkTeamAustralia East20%
CUSTPD-PredictNone
CUSTPD-PrimaveraNone
CUSTPD-Security

Australia East

Australia Southeast

89%

Visual Studio Enterprise Subscription

3b5cbcbf-c2ec-449e-8a22-04a292a4e2d2

None

Visual Studio Enterprise Subscription

8a502b0d-c9e0-427e-b94b-936e5de89cb2

None

In the Subscriptions policy, the settings are currently set to their defaults. These policies govern what level of authorization is required to move subscriptions to and from Azure Active Directory tenants.

With the current default values, it presents the risk that a non-authorized user account could move a subscription to a non-CUST managed tenant (thus removing all management controls).

Figure 3‑4 – Subscription Policies

3.3.3      Cloud Adoption Framework (CAF)

Azure Cloud Adoption Framework recommends organising subscriptions based on both scale and functional requirements, serving as fundamental containers for resources with distinct management, billing, and access control boundaries. Subscriptions should be structured to support organisational requirements for compliance, security, and operational efficiency, while considering service limits and scale constraints. The framework advocates for a hierarchical management approach using management groups for policy inheritance and governance, with subscriptions organised to separate different environments (production, development, test), business units, or workload types. This organisation should align with the enterprise’s billing requirements, administrative model, and security policies while maintaining clear isolation boundaries where needed. Resource management, policy enforcement, and access controls are implemented at the subscription level, making it crucial to plan subscription architecture that accommodates both current needs and future growth while enabling effective cost management and operational oversight.

3.3.4      Recommendation and Assessment

With a new management group structure simplifying logical boundaries, we are freed up to utilise more granular subscription based on one logical boundary (environment) and take advantage of the inheritance that is offered by the hierarchy. This will reduce the complexity and management overhead of managing subscriptions.

Due to the desire to leverage a landing zone per application in the future, it is important to leverage Azure native cost management tools and implement a detailed cost management plan.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Develop a process for resource owners to perform periodic access reviews, policy compliance and cost analysis.521REFACTOR
Develop an CUST cost management plan aligned to the Microsoft ‘managed cloud costs, plan.522REFACTOR
Set controls on subscriptions to prevent them from moving to an unauthorised Entra ID tenant.311REFACTOR

3.4          Resource Groups

3.4.1      Definition

Resource Groups act as logical containers into which Azure resources can be deployed and managed. They act as one of the foundational boundaries for cost, security, and compliance within Azure. Every resource in Azure must belong to a resource group.

3.4.2      Current State

82 Resource Groups have been observed in Azure at the time of publishing this assessment.

There is some inconsistency regarding the naming, capitalisation and logical boundary used to group resources, largely due to the Microsoft default resource groups deployed with some of their solutions e.g. Log Analytics Workspace.

3.4.3      Cloud Adoption Framework (CAF)

Resource Groups are the last logical grouping when it comes to the Azure Hierarchy. Based on the decisions that are made regarding the Management Group and Subscription structures, Resource Groups should be used to group resources that share the same lifecycle or security boundary.

3.4.4      Recommendation and Assessment

Resource Groups should be named relatively simply concentrating on the context of the lifecycle it is grouping, as well as the environment that it belongs to. In the current Azure implementation, it is unclear what the majority of Resource Groups are grouping.

Resource Group Recommendation

–       Investigate and document the lifecycle of Resource Groups.

–       Remove empty resource groups .

–       Review resource groups that have been autogenerated and do not conform to the CUST Naming Standards.

–       Enforce naming standard via a naming module and Azure Policy.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Document the lifecycle of Resource Groups.322REFACTOR
Remove empty Resource Groups.111REFACTOR
Review Resource Groups that do not conform to the CUST naming standards (including autogenerated)211REFACTOR
Review and enforce Naming Standard via naming module and Azure Policy433REFACTOR

3.5          Policy

3.5.1      Definition

Azure Policy is a service that enables the creation, assignment, and management of policies. Policies are used to enforce rules on your resources to ensure they remain complaint to the standards defined by your corporate standards. Azure Policy scans across your resources to identify resources that are not compliant with the implemented policies.

Policies can be assigned at varying scopes, from management groups to resource groups. These offer a method for enabling organization standards from the root management group, with geographic standards applied at their respective management groups.

3.5.2      Current State

Azure Policy is used extensively at CUST with 7262 policies currently in use (mainly compliance Initiatives) and an overall Resource compliance of 90%.

Other policies to be implemented to control the use of:

  • VM SKUs
  • Permitted resource types.
  • Permitted Region deployment

There are several remediation tasks which are pending form the remediation blade in Azure policy

3.5.3      Cloud Adoption Framework (CAF)

The CAF outlines Enterprise Governance as a team sport, requiring all parties to be involved in the process. There is a large amount of effort that comes with translating on-premises based IT policies to the cloud, often requiring the Platform teams and Security to collaborate to translate the policies accurately with the correct language and scope.

Azure Policy within the CAF focuses on four main areas:

  • Cost Management
  • Identity Baseline
  • Security Baseline
  • Resource Consistency

Enabling the use of additional policies is the primary step to ensuring compliance with the CAF, some of the policies recommended:

  • Enable Specific VM SKUs.
  • CIS Audit Initiative
  • Allowed Resource Types (Used to Deny services that have not been CUST ‘endorsed’)

3.5.4      Recommendation and Assessment

Enabling the use of the CIS Audit Initiative will give a grant insight into alignment against the CIS Azure Platform Benchmark. The Policies within the CIS Benchmark align to a great deal of the policies that CUST have currently implemented but are constantly updated by Microsoft with new policies that further align to the CIS benchmark.

Azure Policy can be used to assess compliance of configuration to security standards, defining the expected security standards for the top 10 most used Azure resources and enabling policies to report on the level of compliance would give great insight to the maturity of the Azure environment against the expected documented policies.

Azure Policy Recommendation

–       Create an Azure Policy Governance Framework and enable a team to review, remediate, communicate, and action policies.

–       Define a policy of allowed resource types that have been ‘whitelisted’ for use within the CUST Azure environment.

–       Create a process for onboarding new services and defining appropriate policies to apply to these.

–       Define the use of Azure Policy for the Organisation, who owns and communicates these and how they are life cycled.

–       Consider assigning the Microsoft Defender policies at the Management Group level and unassign them at the subscription level to avoid repeated assignments.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Create an Azure Policy Governance Framework and enable a team to review, remediate, communicate, and action policies.533REFACTOR
Define a policy of allowed resource types that have been ‘whitelisted’ for use within the CUST Azure environment.522REFACTOR
Create a process for onboarding new services and defining appropriate policies to apply to these.533REFACTOR
Consider assigning the Microsoft Defender policies at the Management Group level and unassign them at the subscription level to avoid repeated assignments.411REFACTOR

3.6          Naming and Tagging Standards

3.6.1      Definition

A well-defined naming and tagging standard are a key component to ensuring the ability to quickly locate and manage resources for operational purposes, as well as help to associate usage costs with specific resources and enable a chargeback or show-back model if required. It is important to note that within Azure resource types have different scopes that define the level at which the name must be unique.

3.6.2      Current State

Naming Standards

CUST have implemented a well-structured naming standard based on the Azure resource type, hosting subscription, environment, region and purpose. The Naming standard has been consistently applied except on resources and resource groups that are auto provisioned such as Log Analytics.

Tagging Standards

CUST have implemented a standardised tagging model that are consistently deployed using Azure Policy. The tags that are deployed are a minimal set to identify ownership and could be uplifted to enhance the context of the resource and include operational metadata scenarios that would be required and useful.

It was noted that there are several resources have a tag, but there is no value rendering the tag invalid.

3.6.3      Cloud Adoption Framework (CAF)

Naming Standards

A good naming standard helps to identify resources, be that in cost reporting or in automation scripts. There are two main components that a naming strategy should include, those being Business and Operational.

  • The business-related component to a naming standard should include the organisational information that is required to identity the teams responsible for the resource.
  • The operational side should ensure that names include relevant information that IT teams need, to ensure that the resource can be identified efficiently.

As previously mentioned, CUST have a well-structured naming standard with adequate contextual components to assist in identification of the resource.

Tagging Standards

Tags are essentially a quick way to identify resources via metadata. Tagging is usually very specific to each organisation but the CAF covers a standard examples list of tags that can be used as a template to decide what tags are relevant to your preferred level of scope.

3.6.4      Recommendation and Assessment

The naming convention at CUST should be extended upon to meet the CAF recommendations. This will help with resource scale, identification, and untimely management. Particularly in an environment such has CUST a comprehensive naming standard will assist with management and governance.

There should also be specific tags created for functionality such as automated shut down and startup of virtual machines based on a schedule, but these are specific to virtual machine resources and are an additional capability.

Naming and Tagging Recommendations

–       Review the list of required tags and uplift the tagging standards if required to include additional business and operational context.

–       Re-develop the current tagging policy to enforce the mandatory set of tags, especially on resources.

–       Remediate and update missing tag values.

–       Define and Communicate the Tagging Schema to all Cloud Consumers.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Review the list of required tags and uplift the tagging standards if required to include additional business and operational context.422REFACTOR
Re-develop the current tagging policy to enforce the mandatory set of tags, especially on resources.411REFACTOR
Remediate and update missing tag values.511REFACTOR
Define and Communicate the Tagging Schema to all Cloud Consumers.311REFACTOR

3.7          Cost Management

3.7.1      Definition

Cost Control or Management, much like governance and similar management constructs, depend on a well-managed environment. It is important that the desired level of cost management is factored into the overall design of the Azure Hierarchy. Azure Cost Management gives you the tools to plan for, analyse and reduce your spending to maximise your investment in cloud.

3.7.2      Current State

At present, most of the Azure environment is centrally funded by IT with no show/charge back.

This could result in a cost-blowout as Teams have no ownership for costs and when it comes to provisioning their resources when they start consuming the platform.

In conjunction with a considered tagging schema, cost-show back or cost-chargeback can assist in cost-management.

Motivating resource consumers to:

  • right-size,
  • right-time, and
  • deprovision unused resources.

Creating a culture of Cost Management for all Cloud resource ensures that CUST get value out of the services they build, as a stretch target a cost per product should be aimed for, this can them be linked to a value of the product to the business.

The use of budgets, cost alerts & reserved instances have not been implemented.

As CUST progress to consuming this platform and migrating workloads into Azure, it is important to consider the following cost management strategies from the start to avoid unwanted cost.

  • Right Sizing: Enabling Azure monitor for VM’s to collect usage and thus gain insights into right sizing (less than 5% CPU use).
  • Right Timing: extended use of power cycling in the environment would assist inreducing costs (i.e. turn off when not needed servers).
  • Define ingestion patterns for Logs, presently LAW is centralised to the Sentinel enabled workspace, a potential shift to an operational LAW may be beneficial to support cost management.
  • Ensuring the use of Hybrid licensing for SQL and Windows licenses.

Multiple Orphaned Disks and Network cards shows that resources being deleted are not cleanly done, adding additional costs.

3.7.3      Cloud Adoption Framework (CAF)

Cost Management within the CAF is a large point of discussion as most customers as they move to the cloud are extremely concerned about bill shock or unforeseen costs. Cost Management and Cost Tracking is intended to be accomplished within Azure with the use of Azure Cost Management, Azure Advisor, Tagging, Budgets and Alerting. It is also important to ensure that cost management is taken into consideration from a subscription and management group design point of view as these scopes can enable quick visibility, but also ensure that ability apply the correct permissions for cost management at a higher level.

3.7.4      Recommendation and Assessment

To ensure visibility of cost across Azure, default budgets and alerts should be configured within each subscription or management group to alert budget owners about potential unexpected spend or growth within their environments.

Cost Management Recommendation

–       Configure the Enterprise portal to accurately reflect the organisation structure or cost using departments and accounts.

–       Create a Show Back / Charge Back model for cost management (enable a FinOps Approach).

–       Remove un needed Disks that are not attached to save money (example workbook).

–       Enable diagnostics on VM’s (and PaaS Services) to allow for right sizing recommendations.

–       Enable sensible default budgets and allow teams to create budgets so that costs can be  managed and support review cycles.

–       Create a process to manage and report (and deleted) Orphaned Disks and Nics.

–       Action Azure Advisor insights.

–       Ensure cost definitions are built in to the inception of a services creation .

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Configure the Enterprise portal to accurately reflect the organisation structure or cost using departments and accounts.522REFACTOR

Create a Show Back / Charge Back model for cost management (enable a FinOps Approach).

 

422REFACTOR
Remove un needed Disks that are not attached to save money (example workbook).311REFACTOR
Enable diagnostics on VM’s (and PaaS Services) to allow for right sizing recommendations.411REFACTOR
Enable sensible default budgets and allow teams to create budgets so that costs can be  managed and support review cycles.511REFACTOR

Action Azure Advisor insights.

 

521REFACTOR

Ensure cost definitions are built into the inception of a services creation.

 

422REFACTOR

4            Network

4.1          Overview

Azure has evolved over time to have numerous ways in which networks can be configured for various enterprises. Currently the CAF has two main network topologies that are recommended for implementation, those being, Hub and Spoke and Virtual WAN. As of the start of this month (July 2020), the recommended approach for large enterprise is to leverage the Virtual WAN capability instead of the traditional Hub and Spoke Methodology.

4.2          Virtual WAN

4.2.1      Definition

Azure Virtual WAN is a Microsoft-managed solution where end-to-end global transit connectivity is provided by default. Virtual WAN hubs eliminate the need to manually configure network connectivity. For example, you don’t need to set up or managed user-defined routing (UDR) or network virtual appliances (NVAs) to enable global transit connectivity.

Virtual WAN greatly simplifies the end-to-end network connectivity in Azure and cross-premises, by creating a hub and spoke network architecture that spans multiple Azure regions and on-premises locations.

4.2.2      Current State

CUST have implemented a hub and spoke network design managed by a combined effort of the Network team and the Cloud and DevOps team.

There is a single 50Mbps Express route circuit provided by MegaPort which provides connectivity to the CUST core network. To encrypt Express route, there are 2 VPN’s

  • Australia East – Azure VPN Gateway over express route private peering
  • Australia Southeast – Palo Alto VPN over express route in Australia East.

This network configuration is CUST current approved network topology for Azure connectivity.

Figure 4: Existing Network Topology

4.2.3      Cloud Adoption Framework (CAF)

A network topology based on Azure Virtual WAN is the preferred and recommended enterprise-scale approach for large-scale multi-region deployments where your organisation needs to connect your global locations to both Azure and on-premises locations. A Virtual WAN topology should also be used whenever your organisation intends to use software-defined WAN (SD-WAN) that are integrated with Azure. As a Microsoft-managed service it also reduces the complexity of your network and helps to modernize the landscape.

4.2.4      Recommendation and Assessment

CUST have implemented a well-designed and managed network topology for extending the core network to Azure. This design aligns well to the CAF recommendation for a traditional hub & spoke network topology.

Virtual WAN is a topology for CUST presents an opportunity to utilise the best of both words in terms of unlocking the network landscape within Azure, but also maximizing the investment that has already been made into the existing network.

Microsoft have a guide specifically for the migration of existing global Azure and on-premises footprints to Azure Virtual Hub.

Changing the existing network topology unlocks an opportunity to move away from using networks as a primary security boundary as we did in the previous on-premises worlds, enabling a move towards Zero-Trust networking.

Figure 5: Example of Azure Virtual WAN Topology

The above image depicts a high-level view of a simplistic target state for Azure Virtual WAN.

Benefits:

  • Move away from Network as the primary security boundary (move towards zero – trust)
  • Simplification – Routing intent negates the use of UDR and enables Plug and Play Connectivity for Branches.
  • Alignment to Microsoft Recommendations for large enterprise connected environments.
Networking Topology Recommendation

–       Investigate migrating to a Virtual WAN based Topology to unlock network capability within Azure.

–       Integrate the Azure firewall into Virtual WAN to create a Secured hub core routing capability.

–       Consider using different Express route circuits from other peering locations to ensure redundancy.

–       Configure alerting on Core network services including Express route circuits.

–       Enable connection monitor to monitor connectivity across Express route and VPN.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Investigate migrating to a Virtual WAN based Topology to unlock network capability within Azure.554REIMPLEMENT
Integrate the Azure firewall into Virtual WAN to create a Secured hub core routing capability.523REFACTOR
Consider using different Express route circuits from other peering locations to ensure redundancy.533REFACTOR
Configure alerting on Core network services including Express route circuits.511REFACTOR

Enable connection monitor to monitor connectivity across Express route and VPN.

 

511REFACTOR

4.3          Virtual Networks

4.3.1      Definition

Azure Virtual Networks (VNet) are the fundamental building block for your private network in Azure. VNet enables many types of Azure resources, such as Azure Virtual Machines (VM), to securely communicate with each other, the internet, and on-premises networks. VNet is like a traditional network that you would operate in your own datacentre but brings with its additional benefits of Azure’s infrastructure such as scale, availability, and isolation.

4.3.2      Current State

CUST currently have 21 observable Virtual Networks deployed within Azure.

There is a single Express Route 50mbps throughput and no redundant links.

The VNETS are linked to a hub and spoke model and are not separated by Environment and leverage a common firewall and VNET gateway for traffic flow.

Azure Virtual Network Manager and network groups are used to configure the connectivity and peering from spoke VNET’s to the Hub/Core Azure firewall VNET.

Great use of Virtual Network flow logs in the environment, this has been consistently deployed.

Within the 21 observed VNETs, it was noted that several subnets did not have Network Security Groups (NSGs) associated. Additionally, several subnets lacked route tables. Both conditions require review, consideration, and potential remediation.

List of Subnet without an NSG associated:

VNET NameSubscription IDSubnet Name
vn-sdp-ae-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fAzureBastionSubnet
vn-sdp-ae-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-dsclab
vn-pmv-ae-prd-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-app
vn-pmv-ae-prd-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-sqldb
vn-internet-ae-prd-eslz-001457f2bde-b26b-4950-8c9a-1aab67bf63ebsn-management
vn-internet-ae-prd-eslz-001457f2bde-b26b-4950-8c9a-1aab67bf63ebsn-pe
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-inb
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-ha
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-out
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694AzureBastionSubnet
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-ha
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-inb
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694AzureBastionSubnet
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-out

List of Subnets without route tables associated:

VNET NameSubscription IDSubnet Name
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-out
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-inb
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694AzureBastionSubnet
vn-con-ae-prd-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-ha
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-ha
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694AzureBastionSubnet
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-out
vn-con-as-drp-eslz-001e7d313ee-9482-4278-b7aa-adb964578694sn-dns-inb
vn-internet-ae-prd-eslz-001457f2bde-b26b-4950-8c9a-1aab67bf63ebsn-management
vn-internet-ae-prd-eslz-001457f2bde-b26b-4950-8c9a-1aab67bf63ebsn-pe
vn-pmv-ae-prd-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-sqldb
vn-pmv-ae-prd-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-app
vn-sdp-ae-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fsn-dsclab
vn-sdp-ae-eslz-001f8fc7c8d-53ca-454e-b4d8-45671e7c9b9fAzureBastionSubnet

4.3.3      Cloud Adoption Framework (CAF)

The CAF when aligning to recommendations of Enterprise Scale, recommends the use of capabilities such as:

  • Zero Trust Networking.
  • Network Security Groups use Azure Service Tags rather than IP Addresses.
  • Delegate subnet creation to the Landing Zone / App Resource Group, Owner.

Virtual Networks are the spokes within our Hub and Spoke network topology, even when utilising the Virtual WAN Capabilities.

4.3.4      Recommendation and Assessment

Having an overall Hub-Spoke architecture on the Azure Virtual networks already in is a good starting point. There are some improvements that can be made to bring it into more alignment with the Cloud Adoption Framework for an enterprise environment. These will secure resources and Network flows plus allow for more observability of traffic.

Networking Recommendation

–       Review the deployment of NSGs to all Subnets and remediate where necessary.

–       Review the deployment of route tables to all Subnets and remediate where necessary.

–       Implement a policy to enforce NSG’s and Route tables on all subnets.

–       Uplift Azure Virtual Network Manager and implement Security Admin rules for default platform wide rules.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Review the deployment of NSGs to all Subnets and remediate where necessary.522REFACTOR
Review the deployment of route tables to all Subnets and remediate where necessary.523REFACTOR
Implement a policy to enforce NSG’s and Route tables on all subnets.533REFACTOR
Uplift Azure Virtual Network Manager and implement Security Admin rules for default platform wide rules.532REFACTOR

4.4          Private Endpoints

4.4.1      Definition

Azure Private Endpoints is a network interface that connects you privately and securely to a service powered by Azure Private Link. Private Endpoints uses a private IP address from within your virtual network, effectively bringing the service into your virtual network. This can be used by Azure services such as Azure Storage, Azure Cosmos DB, SQL etc.

4.4.2      Current State

Private Endpoints are currently utilised across the CUST azure network, these are mostly made up of connections to blob endpoints and key vaults, although there is a wide range of other services included.

4.4.3      Cloud Adoption Framework (CAF)

The CAF currently only mentions Service Endpoints, however the same justifications or benefits for the use of Service Endpoints also applies to Private Endpoints.

By utilising Private Endpoints, CUST can effectively pull Azure PaaS resources within our virtual network’s private IP range, when enables secure and private connection to Azure resources. Providing benefits such as:

  • Improved Security for Azure Resources
  • Optimal Routing for Azure Resource traffic from within the Virtual Network
  • Minimal Management Overhead

 

4.4.4      Recommendation and Assessment

Azure Private Endpoints should be enabled where a private facing PaaS resource is required, if there is no Private Endpoint available for that PaaS resource Service Endpoints should be utilised.

Private Endpoint Recommendation

–       Create a Security Pattern to always use Private Endpoints where available and enforce using Azure policy.

–       Use Service Endpoints when Private Endpoints are not available with a routable Service Endpoint subnet, managed by Azure Firewall and Service Endpoint policies attached.

–       Use Azure policy to create and lifecycle manage DNS records associated with Private endpoints.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Create a Security Pattern to always use Private Endpoints where available and enforce using Azure policy.422REFACTOR
Use Azure policy to create and lifecycle manage DNS records associated with Private endpoints.522REFACTOR

4.5          Azure Firewall

4.5.1      Definition

Azure Firewall is a managed, cloud-based network security service that protects your Azure Virtual Network resources. It is a fully stateful firewall as a service with built-in high availability and unrestricted cloud scalability.

4.5.2      Current State

CUST physical network connections are configured via express rout circuits with an Azure VPN overlay in Australia East and Palo Alto VPN for Australia Southeast using the same Express route circuit.

All traffic is routed through the Azure Firewalls deployed in the Hub VNET and routing enforced using a route table where it exists.

Whilst this is a legacy Hub and Spoke pattern, it is still effective but could certainly be optimised using the Virtual WAN Secure Hubs as mentioned earlier in the document.

4.5.3      Cloud Adoption Framework (CAF)

The CAF largely refers to the use of Azure Firewall via design recommendations:

  • Azure Firewall should be deployed in the Virtual WAN Hubs for east-west and/or north-south traffic protection / filtering.
  • Within a Hub and Spoke topology Azure Firewall should be deployed within the Central HUB for east-west and/or north-south traffic protection / filtering.
  • Create global Azure Firewall Policy to govern security posture across the global network environment and assign it to tall Azure Firewall instances.
  • Enable the use of Azure Service Tags where possible instead of using dedicated IP Addresses.

4.5.4      Recommendation and Assessment

CUST have implemented a well-designed and managed network topology for extending the on-premises network to Azure. This design aligns well to the CAF recommendation for a traditional hub & spoke network topology but could be optimised to leverage Azure Virtual WAN.

Azure Firewall Recommendation

–       Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs.

–       Uplift the firewall policy and create a Hierarchy, set a Global policy to all existing firewalls that contains ‘must have’ rules and use local policies for individual and localised rules.

–       Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs.

–       Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny.

–       Configure Azure firewall IDS in Deny mode.

–       Consider using IP groups  to reduce IP table rules usage.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs.543REIMPLEMENT
Uplift the firewall policy and create a Hierarchy, set a Global policy to all existing firewalls that contains ‘must have’ rules and use local policies for individual and localised rules.422REFACTOR
Review the current network Architecture and consider re-implementing onto Azure Firewalls deployed into Virtual Hubs.543REIMPLEMENT
Configure Azure firewall policy to enforce threat intelligence mode to Alert and Deny.532REFACTOR
Configure Azure firewall IDS in Deny mode.522REFACTOR
Consider using IP groups  to reduce IP table rules usage.521REFACTOR

5            Monitoring & Logging

5.1          Overview

Monitoring within Azure is completed by four services; Azure Monitor, Azure Service Health, Azure Advisor and Microsoft Defender for Cloud, all of which serve a different purpose or function. Within this section we will mainly be looking at Azure Monitor as a service and its capabilities as a single unified hub for all monitoring and diagnostics data within Azure. To maximise the value out of Monitoring we need to ensure that we have a well-defined and designed Logging platform, or as the CAF calls it Monitoring Data Platform.

Figure 5‑1: Azure Monitor Service Overview

5.2          Logging

5.2.1      Definition

Logging within Microsoft Azure is comprised of three different classifications of what is referred to as “Platform Logs”. These logs are used to provide detailed information ranging from diagnostic information about specific resources to auditing information for activity logs or Azure Active Directory.

Resource Logs

Resource Logs, previously known as “Diagnostic Settings”, provide insight into operations that were performed within an Azure resource (the data plane), for example getting a secret from a Key Vault, or making a request to a database. The content of resource logs varies by the Azure service and resource type.

Activity Logs

Activity Logs, provides insight into the operations on each Azure resource within the context of a subscription, from management plane, in addition to updates on Service Health events. Activity logs can be used to determine what, who and when for any write operations (PUT, POST, DELETE) taken on the resources within your subscription.

Microsoft Entra Logs

Entra Logs, contains all the sign-in activity and audit trail of changes made in the Entra for a particular tenant.

5.2.2      Current State

CUST currently have recently implemented standards around the capturing of log data within Azure. Logging capture has been implemented into a single Sentinel enabled LAW.

It was found, however, that there are several ‘default’ workspaces that are not within the logging construct and do not follow naming conventions within the CUST environment.

5.2.3      Cloud Adoption Framework (CAF)

The CAF has three strategies that can be applied to Logging, those being:

  • CentralisedAll logs are stored in a central workspace and managed by a single team. With Azure Monitor providing differential access on a per-team basis. In this case, it is easy to manage, search across resource and cross-correlate logs. But one of the main limitations of this model is the administrative overheard that eventually comes as the use of the workspace grows. This model is also known as the Hub and Spoke.
  • DecentralisedEach team has their own workspace created in a resource group that they can each own and manage, where log data is segregated per resource. In this case, the workspace can be kept security and access control is able to be kept consistent with resource access. But it’s extremely difficult to cross-corelate logs. Users and Enterprises who need to create a broad view of the Azure landscape will find it difficult to analyse the data in a meaningful way without large amounts of effort.
  • HybridIt is common to see organisations attempt to deploy both above strategies in parallel which leads to a complex, expensive, and hard-to-maintain configuration that also has gaps in coverage.

5.2.4      Recommendations and Assessment

CUST should look to implement a centralised data collection / logging capability within Azure. This does not mean a singular instance of Log Analytics but keeping the number as small as possible will ensure that the environment is simple and manageable.

Whilst keeping the number of Log Analytics Workspaces to a minimum it is important to note that ingesting data across regions within Azure does come with a cost. So where required a separate Log Analytics Workspace should be created per geographical region.

Figure 5‑2: Diagnostic Settings Logging Pattern

Within the context of each Log Analytics resource there is also multiple levels of access controls that can be applied. As there has already been a large amount of upfront effort within the management of access to resources within Azure, utilising the Azure Permissions access model is recommended, this enables low operational overhead whilst enabling re-use of granular permissions declared against the Azure resources.

Logging Recommendations

–       Consider segregating Operational logging data from Security logging data to save on cost.

–       Establish Logging patterns for an all services (PaaS and IaaS) and document them and deploy consistently via Azure Policy.

–       Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data.

–       Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards.

–       Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Consider segregating Operational logging data from Security logging data to save on cost.533REFACTOR
Establish Logging patterns for an all services (PaaS and IaaS) and document them and deploy consistently via Azure Policy.542REIMPLEMENT
Create a Service activation approach (patterns) that enable creation and consumption of diagnostics logging data.542REIMPLEMENT
Define metrics and monitoring standards for services deployed within Azure align them to business operational requirements and standards.533REIMPLEMENT
Develop an operational framework to enable support and run teams to effectively manage hosting infrastructure.522REIMPLEMENT

5.3          Monitoring

5.3.1      Definition

Azure Monitor is the native Azure service to enable the maximum availability and performance potential of your applications and services by enabling a single platform for collecting, analysing, and acting on telemetry from your cloud and on-premises environments.

Azure Monitor is targeted to help you proactively identify issues within your applications and infrastructure whilst also providing insight into the performance of these services.

In August 2018, Microsoft consolidated Log Analytics, and Application Insights into the service we now know as Azure Monitor. This was done to enable a single integrated experience of monitoring Azure resources and hybrid environments.

5.3.2      Current State

CUST have implemented Azure monitoring but there is a limited capacity.

Currently, there has been no alerting configured to monitor Azure usage and events from logs captured within Log Analytics. Application Insights has been configured on 2 instances to gain low level application telemetry that is exposed into the Azure Monitor platform. Azure Dashboards are not used to assist with monitoring applications and network utilisation.

5.3.3      Cloud Adoption Framework (CAF)

Azure Monitor is the Azure native platform service that provides a single source for monitoring Azure resources. It can monitor all layers of the stack, starting with Tenant services, such as Azure Active Directory Services through to Application Insights such as SQL Queries between an Azure Web Application and an Azure SQL Database. It is because of this full stack capability that the CAF recommends utilising Azure Monitor wherever possible for monitor as opposed to 3rd party tooling. The CAF does cover Alerting recommendations for within Azure Monitor, but this is more just what you can alert on with Azure Monitor rather than best practices.

5.3.4      Recommendations and Assessment

We recommend that CUST have a documented Monitoring strategy that takes a service-orientated approach to clearly define what monitoring is to be used for the infrastructure layer, resource, scope, and the method that is to be used for monitoring. The below recommendations are what should be implemented in Azure Monitor and, any application monitoring requires definition and validation of effectiveness to ensure the end-to-end implementation meets the business requirements.

Monitoring Recommendations

–       Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone.

–       Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example), alert on anomalies.

–       Enable VM Insights and performance counters for relevant workloads.

–       Configure Service and Resource health alerts and notify central platform and operational teams when issues occur.

–       Define and implement Monitoring for all PaaS Services.

–       Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications.

–       Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Develop and document a Monitoring strategy for Azure resources, include an alerting framework for resources and consider a default set of alerts in every landing zone.532REIMPLEMENT
Enable Network Monitoring and Endpoint monitoring (ExpressRoute to Datacentre for example).522REFACTOR
Configure Service and Resource health alerts and notify central platform and operational teams when issues occur.421REFACTOR
Define and implement Monitoring for all PaaS Services.322REFACTOR
Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications.322REFACTOR
Document reference Azure architecture & blueprint artefacts for high priority / cloud-based hosting patterns.421REFACTOR

5.4          Azure Dashboards

5.4.1      Definition

Azure Dashboards are a focused and organized view of cloud resources within the Azure Portal. Dashboards can be used as a workspace where you can quickly launch relevant tasks for your day-to-day operational activities, such as the main resource group for the application you are currently working on. Dashboards are also able to be published and shared across an Azure Tenant, to enable a focused visibility on specific environments or capabilities within Azure or even On-Premises data.

5.4.2      Current State

CUST have not implemented the use of dashboards to monitor the usage and health of any resources running on Azure. Dashboards can be created to monitor the express route utilisation and cover the standard infrastructure landscape that support the Platform to give a snapshot view of the operational state of core services.

5.4.3      Cloud Adoption Framework (CAF)

Dashboarding is not specifically mentioned within the CAF.

5.4.4      Recommendations and Assessment

Azure Dashboards are the best way to get visibility across an applications environment. As the data to make use of Azure Dashboard will already be collected there is minimal effort required to configure the use of Azure Dashboards across CUST.

Dashboard Recommendations

–       Create the ability for Templated Dashboards / Workbooks as shareable entities.

–       Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications.

–       Extend the use of Dashboards / workbooks for Infrastructure monitoring.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Create the ability for Templated Dashboards / Workbooks as shareable entities.211REFACTOR
Extend the use of Dashboards / workbooks for Production Environments of Azure hosted Applications.211REFACTOR
Extend the use of Dashboards / workbooks for Infrastructure monitoring.321REIMPLEMENT

6            Identity

6.1          Overview

In any environment, whether on-premises, hybrid, or cloud-only, IT needs to control which administrators, users, and groups have access to resources. Identity and access management (IAM) services enable you to manage access control in the cloud.

Several options are available for managing identity in a cloud environment. These options vary in cost and complexity. A key factor in structuring your cloud-based identity services is the level of integration required with your existing on-premises identity infrastructure.

Figure 6‑1: Azure Active Directory complexity scale

6.2          Entra ID

6.2.1      Definition

Entra ID is Microsoft’s cloud-based identity and access management in Azure Cloud. It’s at the centre of authentication and authorisation for users access as well as applications and machine identities.

Entra ID uses modern authentication protocols and is fully accessible via Microsoft Graph API for programmatic accesses. Depending on the licencing model chosen, a range of features are available to enhance visibility and security controls.

A common scenario for businesses is to use Entra ID in a hybrid configuration. Apart from startups that would be Cloud native, most companies would have an existing identity solution on-premises to manage their users’ accesses to internal and external resources.

6.2.2      Current State

CUST have a single Entra ID tenant within Azure.  Microsoft Entra Connect with Connect Sync has been configured to synchronise identity with on-premises domain.

Tenant Display NameDomain NamesTenant IDSubscriptions

CUST PTY LTD

 

CUST CUSTPTYLTD.onmicrosoft.comaa391e1f-5f1c-48e3-aef0-4c1c91da1718Yes

The current secure Score on Entra is 74% and the following areas of improvement should be targeted, targeting the Low Implementation cost and Low user impact first.

NameScore ImpactCurrent ScoreMax ScoreUser ImpactImplementation CostStatus
Stop clear text credentials exposure1.8305LowLowTo address
Modify unsecure Kerberos delegations to prevent impersonation1.8305LowLowTo address
Reduce lateral movement path risk to sensitive entities1.8305LowLowTo address
Protect and manage local admin passwords with Microsoft LAPS1.8305LowLowTo address
Resolve unsecure account attributes1.8305LowLowTo address
Stop weak cipher usage1.8305LowLowTo address
Change password of built-in domain Administrator account1.8305LowLowTo address

User settings for App registrations and Administration portal are appropriate. Linkedin account connections is enabled for all users, allowing this grants access to user’s properties, confirm that this is required or not.

External collaboration settings are configured appropriately, other than the Collaboration restrictions which may need review.

CUST are leveraging named locations

6.2.3      Cloud Adoption Framework (CAF)

For organizations with existing on-premises Active Directory infrastructure, directory synchronization is often the best solution for preserving existing user and access management while providing the required IAM capabilities for managing cloud resources. This process continuously replicates directory information between Entra ID and on-premises directory services, allowing common credentials for users and a consistent identity, role, and permission system across your entire organisation.

Directory synchronization assumptions: Using a synchronized identity solution (Entra Connect) assumes the following:

  • You need to maintain a common set of user accounts and groups across your cloud and on-premises IT infrastructure.
  • Your on-premises identity services support replication with Entra ID.

Cloud-hosted domain services assumptions: Performing a directory migration assumes the following:

  • Your workloads depend on claims-based authentication using protocols like Kerberos or NTLM.
  • Your workload virtual machines need to be domain-joined for management or application of Active Directory group policy purposes.

Entra federation services assumptions: Using EFS assumes the following:

  • Your workloads require a single sign-on capability cross multiple domains within an organisation.

6.2.4      Recommendations and Assessment

CUST have a good implementation of Entra ID and are using Azure managed identities as a principal type within the platforms DevOps pipelines.

Dashboard Recommendations

–       Review Linkedin account connections being enabled for all users.

–       Allow invitations only to the specified domains, ensure there is a regular review of the targeted domains.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Review Linkedin account connections being enabled for all users.311REFACTOR
Allow invitations only to the specified domains, ensure there is a regular review of the targeted domains.311REFACTOR

6.3          Role Based Access Control

6.3.1      Definition

Access management for cloud resources is a critical function for any organization that is using the cloud. Azure role-based access control (Azure RBAC) helps you manage who has access to Azure resources, what they can do with those resources, and what areas they have access to.

Azure RBAC is an authorization system built on Azure Resource Manager that provides fine-grained access management of Azure resources.

6.3.2      Current State

CUST have implemented RBAC roles for user access to subscription resources and Entra objects. There is a consistent implementation of RBAC roles has been observed, however, they are generic and aligned to the Azure Roles instead of an CUST Persona/Operating model within the platform and generated landing zones.

It was also noted that CUST are using Access Packages to allow Teams to request access to the relevant resources, whilst good practice the requestable roles should be Persona/Role aligned and protected with Azure ABAC to streamline the approval process and ensure platform users are not over permissioned.

6.3.3      Cloud Adoption Framework (CAF)

Only Grant the access users need:

Using Azure RBAC, you can segregate duties within your team and grant only the amount of access to users that they need to perform their jobs. Instead of giving everybody unrestricted permissions in your Azure subscription or resources, you can allow only certain actions at a particular scope.

When planning your access control strategy, it’s a best practice to grant users the least privilege to get their work done. Avoid assigning broader roles at broader scopes even if it initially seems more convenient to do so. When creating custom roles, only include the permissions users need. By limiting roles and scopes, you limit what resources are at risk if the security principal is ever compromised.

The following diagram shows a suggested pattern for using Azure RBAC.

Figure 8: Role Based Access Control Pattern

Limit the number of subscription owners:

You should have a maximum of 3 subscription owners to reduce the potential for breach by a compromised owner. This recommendation can be monitored in Microsoft Defender for Cloud.

Use Entra ID Privileged Identity Management:

To protect privileged accounts from malicious cyber-attacks, you can use Entra ID Privileged Identity Management (PIM) to lower the exposure time of privileges and increase your visibility into their use through reports and alerts. PIM helps protect privileged accounts by providing just-in-time privileged access to Entra ID and Azure resources. Access can be time bound after which privileges are revoked automatically.

6.3.4      Recommendations and Assessment

CUST have implemented a good RBAC model for the Entra tenants. Continual review of the RBAC model should be implemented to assess the organization access requirement and ensure a least privileged model is implemented to reduce the impact and risk of compromised or mishandled identity access.

Role Based Access Recommendations

–       Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments and access packages.

–       Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones.

–       Consider establishing a cloud-based remote-access solution for Azure Administration.

–       Re define PIM to audit and validate that PIM escalation are tightly controlled.

–       Conduct PIM access reviews periodically.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Redefine the use of Roles within Azure in adherence with the least privilege model and adopt a ‘Persona’ driven approach to role assignments and access packages.544REIMPLEMENT
Consider the use of Azure ABAC conditions for privileged roles within the platform and landing zones.533REFACTOR
Consider establishing a cloud-based remote-access solution for Azure Administration.544REIMPLEMENT
Re define PIM to audit and validate that PIM escalation are tightly controlled.422REFACTOR

6.4          Conditional Access

6.4.1      Definition

Conditional Access is the tool used by Azure Active Directory to bring signals together, to make decisions, and enforce organizational policies. Conditional Access is at the heart of the new identity driven control plane.

By using Conditional Access policies, you can apply the right access controls when needed to keep your organization secure and stay out of your user’s way when not needed.

Figure 9: Conditional Access Pattern

6.4.2      Current State

CUST currently have 28 policies enabled. Policies configured within the CUST Entra Tenant are aligned with ASD’s Blueprint for Secure cloud. At a high-level, these policies require the following standards:

Base Protection Policy

  • All users and administrators required to use MFA
  • Blocks legacy authentication
  • Requires compliant or hybrid Entra joined devices
  • Blocks access from non-trusted locations

Privileged Access Policy

  • Requires MFA for all privileged role holders
  • Additional device compliance requirements for privileged accounts
  • Session controls including sign-in frequency
  • Restricted access to administration portals
  • Blocks access from non-corporate devices

Session Controls

  • Sign-in frequency: 4 hours
  • Persistent browser session: Disabled
  • Continuous access evaluation: Enabled
  • Device filter: Compliant devices only
  • Application enforced restrictions for sensitive data

6.4.3      Cloud Adoption Framework (CAF)

Conditional Access is not specifically mentioned within the CAF.

The CAF references using Security defaults as a baseline. Security defaults make it easier to help protect your organization from these attacks with preconfigured security settings:

  • Requiring all users to register for Microsoft Entra Multi-Factor Authentication.
  • Requiring administrators to perform multi-factor authentication.
  • Blocking legacy authentication protocols.
  • Requiring users to perform multi-factor authentication when necessary.
  • Protecting privileged activities like access to the Azure portal.

6.4.4      Recommendations and Assessment

Conditional Access Policies are configured appropriately within the CUST tenant.

6.5          Privileged Identity Management

6.5.1      Definition

Privileged Identity Management (PIM) is a service in Microsoft Entra ID that enables you to manage, control, and monitor access to important resources in your organization. These resources include resources in Entra ID, Azure, and other Microsoft Online Services such as Microsoft 365 or Microsoft Intune

Privileged Identity Management provides time-based and approval-based role activation to mitigate the risks of excessive, unnecessary, or misused access permissions on resources that you care about. Features of Privileged Identity Management:

  • Provide just-in-timeprivileged access to Entra ID and Azure resources
  • Assign time-boundaccess to resources using start and end dates
  • Require approvalto activate privileged roles
  • Enforce multi-factor authenticationto activate any role
  • Use justificationto understand why users activate
  • Get notificationswhen privileged roles are activated
  • Conduct access reviewsto ensure users still need roles
  • Download audit historyfor internal or external audit

6.5.2      Current State

PIM has been implemented across CUST Entra tenant. A Consistent implementation has been observed for privileged Entra roles.

There were 2 assignments which were listed as permanent within the Global Administrator role which seem to be break glass accounts, although they will need validation.

6.5.3      Cloud Adoption Framework (CAF)

Microsoft recommends the following Azure roles managed by Privileged Identity Management as a baseline

  1. Global administrator
  2. Security administrator
  3. User administrator
  4. Exchange administrator
  5. SharePoint administrator
  6. Intune administrator
  7. Security reader
  8. Service administrator
  9. Billing administrator
  10. Skype for Business administrator

Microsoft recommends that you manage Owner roles and User Access Administrator roles of all subscriptions/resources using Privileged Identity Management.

Microsoft recommends you manage all roles with guest users using Privileged Identity Management to reduce risk associated with compromised guest user accounts.

Microsoft recommends that you bring Entra ID role-assignable groups under management by Privileged Identity Management.

Microsoft recommends you have zero permanently active assignments for both Entra ID roles and Azure roles other than the recommended two break-glass emergency access accounts.

Microsoft recommends you set up Azure log monitoring to archive audit events in an Azure storage account for greater security and compliance.

6.5.4      Recommendations and Assessment

Privileged Identity Management Recommendations

–       Review permanent Global Administrator assignments.

–       Review PIM role assignments against Microsoft PIM baseline.

–       Conduct regular review of PIM alerts.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Review permanent Global Administrator assignments.311REFACTOR
Review PIM role assignments against Microsoft PIM baseline.311REFACTOR
Conduct regular review of PIM alerts.411REFACTOR

7            Security

7.1          Defender for Cloud

7.1.1      Definition

Microsoft Defender for Cloud is a unified infrastructure security management system that strengthens the cloud security posture and provides advanced threat protection across your hybrid workloads in the cloud – whether they are in Azure or not – as well as on-premises. Microsoft Defender for Cloud comes with features built-in such as:

  • Workflow Automation: backed by Azure Logic Apps which allow you to automate responses to certain alerts e.g., send email notifications and raise a ticket in Service Now.
  • Azure Policies and Compliance: Visibility and real-time reporting on your security posture according to predefined or custom Azure Policies.
  • Advanced Cloud Defence:
    • Just-in-time virtual machine access
      • Opens public IP to allow for RDP/SSH connection (Updates NSG on VM)
      • Disable recommendations around JIT access as it contradicts CIS and Public endpoints
    • Adaptive application controls
      • Automated intelligence framework that helps reduce attack service to VMs
      • Create alerts if unknown or trusted applications run on VM that has not been whitelisted.
    • Threat Protection: provides comprehensive defence for your environment such as Azure compute resources (VMs, App Services, Containers), data resources (Azure SQL, CosmosDB) and service layers (VNet, Key Vaults).

7.1.2      Current State

Microsoft Defender for Cloud is broken up into 4 main areas of management those being:

  • Policy & Compliance
  • Resource Security Hygiene
  • Threat Protection
  • Advanced Cloud Defence

As it currently stands CUST are consuming the Azure Defender CSPM plan in partial workload protection configuration in all the onboarded subscriptions. There is a consistent configuration across all the subscriptions and plans with a good coverage of resources.

The posture within Policy and Compliance is based on the Microsoft “Secure Score” which is calculated based on the ratio between your healthy resources and your total resources. A “Healthy” resource is determined by any recommendations that are generated from the Microsoft Defender for Cloud. CUST’s current Secure Score across its subscriptions are at the time of publishing this assessment:

Threat protection provides insight into perceived threats within your azure environments specifically alerting on issues that are against Azure Compute, Data or Service layer resources. These alerts are usually paired with recommended remediation steps but in some instances can also be used to trigger downstream actions.

There is 1 active alert were observed within Threat Protection

There are several security recommendations noted within Defender, some labelled as critical and high which require intervention.

7.1.3      Cloud Adoption Framework (CAF)

The Cloud Adoption Framework largely focuses on the use of Defender for cloud to pre-emptively detect vulnerabilities across the Azure environment. The recommendations from the CAF are directly aligned to the recommendations that can be seen within the Defender for Cloud as it looks to assess the deployed resources within your Azure environments. This includes things such as Just-in-time (JIT) access implementation for Azure Virtual Machines.

7.1.4      Recommendations and Assessment

CUST have done a comprehensive job collecting and feeding data Defender for Cloud by enabling data collection for Azure resources to Log Analytics. To take advantage of this effort, a review & action of the recommendations active in Defender to further secure azure resources where appropriate.

Defender for Cloud Recommendations

–       Investigate dismissing false positive recommendations within Defender.

–       Establish a weekly review of Defender recommendations.

–       Review the Regulatory Compliance reporting within Defender and consider the recommended controls, some of which require Azure policy enforcement or configuration.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Investigate dismissing false positive recommendations within Defender.411REFACTOR
Establish a weekly review of Defender recommendations.421REFACTOR
Review the Regulatory Compliance reporting within Defender and consider the recommended controls, some of which require Azure policy enforcement or configuration.422REFACTOR

7.2          Azure Sentinel

7.2.1      Definition

Microsoft Sentinel is a scalable, cloud-native, Security Information Event Management – SIEM. Azure Sentinel delivers intelligent security analytics and threat intelligence across the enterprise, providing a single solution for alert detection, threat visibility, proactive hunting, and threat response.

Sentinel provides a correlation point for alerts and information from all Microsoft services.

7.2.2      Current State

CUST currently leverage Azure Sentinel out of the central Log Analytics workspace.

There are several watchlists in place but could be extended to detect sensitive accounts & break-glass account compromise.

7.2.3      Cloud Adoption Framework (CAF)

Whilst there are no CAF specific recommendations, Microsoft’s guidance is that CUST should define and document clear security objectives and asset identification, followed by establishing governance through RBAC and retention policies. Security operations benefit from automated responses and cross-workspace monitoring capabilities. The management approach requires regular cost monitoring, alert tuning, and archival strategies. Operational effectiveness depends on logical resource organisation and health monitoring of data sources. These practices work together to create a comprehensive security information and event management (SIEM) solution that supports both security and business objectives.

7.2.4      Recommendations and Assessment

Azure Sentinel has an extremely easy integration model to Azure services, Enabling some Sentinel’s default capabilities would provide additional security insight to potential security issues within CUST’s Azure environment.

Sentinel Recommendations

–       Enable UEBA to detect and Identify Insider threats in CUST’s environment and their potential impact.

–       Review and extend the use of Microsoft Sentinel watchlists to investigate threats, import data, reduce alert fatigue & enrich event data derived from multiple data sources.

–       Refer to the Microsoft community Playbooks on Github for other security-related playbooks as per business requirements.

–       Re-validate enabled \ disabled playbooks for response to incidents.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Enable UEBA to detect and Identify Insider threats in CUST’s environment and their potential impact.412REFACTOR
Review and extend the use of Microsoft Sentinel watchlists to investigate threats, import data, reduce alert fatigue & enrich event data derived from multiple data sources.421REFACTOR
Refer to the Microsoft community Playbooks on Github for other security-related playbooks as per business requirements.312REFACTOR
Re-validate enabled \ disabled playbooks for response to incidents.422REFACTOR

8            DevOps

8.1          Definition

DevOps is the union of people, processes, and technology to continually provide value to customers. DevOps enables formerly siloed roles: development, IT operations, quality engineering and security to coordinate and collaborate to produce better, more reliable products.

By adopting a DevOps culture along with DevOps practices and tools, teams gain the ability to better respond to customer needs, increase confidence in the applications they build and achieve business goals faster.

8.2          People Process and Culture

8.2.1      Definition

The Cloud Adoption Framework identifies Devops only in terms of building a Cloud Adoption Plan.

You can’t buy DevOps, as DevOps is not a software, tool, process, company, or person, it’s a methodology used especially by IT professionals.

DevOps is the correlation of people, process, and products to enable continuous delivery of value to end users. The outcomes are tightly connected to allow for frequent releases and at the same time to keep the same level of quality.

People

Initially, stakeholders would need to be identified, ensuring that everybody delivering value to the business are working tightly together on the common goal of adding value to the customer. The latest study by Gartner, and from this study we can deduce that the biggest concern is ‘People’, with process and products being deemed less critical. From this, we can assume that having highly motivated people with good collaboration is necessary.

Process

Next – DevOps is about improving process because even if you have highly motivated people that are working well together, you may still have several business processes which may get in the way, and this could really block innovation. For instance, having to seek approval from long chain advisory boards before implementing changes or being restricted to doing things in a certain way can impede innovation. The process of designing, building, and testing software should be well presented to each individual team member, making them aware of all parts of the development process. The implementation of DevOps can be hard work, as it completely changes the company’s structure.

The core of this is enabling efficient flow of collaboration and making sure that business structure and processes do not get in the way but instead have processes and practices that help improve the value and delivery to your customers.

Products

Products, tools, and services that can help enable different DevOps practices and different teams which can be used to make things easier. From a very high level, these tools include Microsoft Azure, which offers a lot of different products and services.

Figure 11: Sample Mapping of skills to IT roles in a cloud-hosted environment.

The following table represents the suggested RACI per Team but also how each team will interact with one another.

TeamSolution deliveryBusiness alignmentChange managementSolution operationsGovernancePlatform operationsPlatform automation
Strategy teamCAACCII
Adoption teamACRCIII
Ops teamCCRACAC
Gov teamCIICARI
Aligned CapabilityCloud adoptionCloud strategyCloud strategyCloud operationsCCoE & GovernanceCCoE and Cloud platformCCoE and Cloud automation
          

8.2.2      Current State

CUST have a dedicated DevOps team that works closely alongside key Infrastructure and Security teams to deliver platforms and solution.

CUST currently leverage various teams across the business that are autonomous from each other but still collaborate on joint project where required. Responsibility for Azure infrastructure mostly falls to the cloud Infrastructure and Security chapters.

Within the DevOps team, there are highly skilled engineers but it is not clear what functions they perform within the team and platform. There is clear deliniation between Platform and Landing zones but there is no clear definition that same structure in relation to personas and functions so it is not clear to what degree onboarding teams will be enabled and autonomous.

There does not appear to be a defined Enterprise approach to DevOps and architecture so the strategy is loosley defined which may result in pcokets of high and low skill areas within the organistaion. Sprints within the DevOps team occur monthly or every 4 weeks and there are no retrospectives captured to understand if adequate progress has been made against the backlog of work.

8.2.3      Cloud Adoption Framework (CAF)

The Azure CAF generally talks to the uptake of services in this space and how organisations can flex and change to allow for this.

  • Digital estate rationalisation:What are the top 10 priority workloads in the adoption plan? How many additional workloads are likely to be in the plan? How many assets are being considered as candidates for cloud adoption? Are the initial efforts focused more on migration or innovation activities?
  • Organisation alignment:Who will do the technical work in the adoption plan? Who is accountable for adherence to governance and compliance requirements?
  • Skills readiness:How many people are allocated to perform the required tasks? How well are their skills aligned to cloud adoption efforts? Are partners aligned to support the technical implementation?

8.2.4      Recommendations

People, Process, Technology Recommendations

–       Establish and document enterprise DevOps standards.

–       Define the RACI matrix for current teams to identify and address any existing silos.

–       Create a CoE forum/framework to facilitate team collaboration without management or process barriers, promoting shared ownership.

–       Use the forum to iteratively enhance services with transparent feedback for seamless improvements and adoption.

–       Centralise ownership of targeted IaC components/modules and implement a feedback loop for overall improvements (e.g., distributed commit, central creation).

–       Formalise the operational support approach and align capabilities to ensure operational resilience.

–       Reduce sprint time to 2 week sprints and establish a process to perform restrospectives on delivered sprint to capture learnings and imrpove.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Establish and document enterprise DevOps standards.535REIMPLEMENT
Define the RACI matrix for current teams to identify and address any existing silos.532REIMPLEMENT
Create a CoE forum/framework to facilitate team collaboration without management or process barriers, promoting shared ownership.532REIMPLEMENT
Use the forum to iteratively enhance services with transparent feedback for seamless improvements and adoption.522REIMPLEMENT
Centralise ownership of targeted IaC components/modules and implement a feedback loop for overall improvements (e.g., distributed commit, central creation).533REIMPLEMENT
Formalise the operational support approach and align capabilities to ensure operational resilience.543REIMPLEMENT
Reduce sprint time to 2 week sprints and establish a process to perform restrospectives on delivered sprint to capture learnings and imrpove.521REIMPLEMENT

8.3          Architecture

8.3.1      Definition

The following section includes a list of specific Architecture artefacts to be created/modified to include automation / DevOps capabilities and their associate interactions across the enterprise.

These artefacts will be critical in driving rationalisation, & integration objectives.

Strategy & Vision

The intent of the DevOps and Automation strategy is to define the guiding principles and characteristics of the associated roadmaps, blueprints, and reference architectures. The architecture team should utilise the governance structures defined in later sections, along with direct business and wider chapters within CUST.

Enterprise Architecture

To support the reuse, consolidation and efficiency potential of DevOps, there must be a targeted effort in defining the business and technical interfaces.

In order achieve the goal of repeatable end to end build pipelines of business applications, the storage, sharing, and control of application code, data and configuration information should be defined up front and not as a by-product of early DevOps / Cloud adoption efforts.

Application Architecture

The application architecture of business applications should be defined before or in the very early stages of DevOps onboarding or cloud migration / build. Clear guidance of the intended architecture will influence hosting, build, validation, and branching strategies of not just the application in scope but it’s child, parent, and other applications in interacts / shares with. Clearly articulating the application architecture will allow Infrastructure and Operations teams to better shape the establishment of the Services and Systems for consumption by applications teams. Over time, as the hosting reference architecture and service catalogues of the new contemporary platforms matures it will in part begin to influence the application architecture in bidirectional way.

Reference Architecture and Blueprints

A series of reference architecture & blueprint artefacts should be completed for high priority / cloud-based hosting patterns. This will allow the development of deployment pipelines with the necessary building blocks end to end build pipelines.

8.3.2      Current State

From a Platform perspective, the Architectural and design documentation does not describe the assessed environment accurately. Within the design documenation that was assessed, there is limited detail and an unsutable level of context which in turn makes it difficult the understand for Architecture, and difficult to build too for an engineering team. These discrepincies generally lead to a substandard output and an increase in technical debt and refactoring post build. This mis-alignment makes it challenging to communicate the current state to new Team members or to share with outside teams that will be consumers of this platform, placing greater pressure and reliance on the Cloud DevOps team.

There is currently no defined process for defining architectural planning workflows or documentation of CUST’s Azure tenant or any solutions deployed to it. During the discovery workshop, it was described as mostly an ad-hoc depending on the solution and the chapter that oversees it.

8.3.3      Cloud Adoption Framework (CAF)

The Cloud Adoption Framework enterprise-scale landing zone architecture represents the strategic design path and target technical state for an organization’s Azure environment. It will continue to evolve alongside the Azure platform and is defined by the various design decisions an organization must make to map an Azure journey.

Not all enterprises adopt Azure the same way, so the Cloud Adoption Framework enterprise-scale landing zone architecture varies between customers. The technical considerations and design recommendations in this guide might yield different trade-offs based on your organization’s scenario. Some variation is expected, but if you follow the core recommendations, the resulting target architecture will set your organization on a path to sustainable scale.Cloud foundations workshop, it was described as mostly an ad-hoc depending on the solution and the chapter that oversees it.

8.3.4      Architecture Reccomendations

Architecture Recommendations

–       Document Architecture Strategy and Vision as set of guiding principles / technical charter.

–       Identify and document Information Architecture to the treatment of data sensitivity.

–       Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables).

–       Uplift current Platform Architectural documentation to accurately reflect the implemented platform.

–       Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework.

–       Create solutions module repository of common reference Terraform modules for deploying Azure resources.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Document Architecture Strategy and Vision as set of guiding principles / technical charter.533REIMPLEMENT
Identify and document Information Architecture to the treatment of data sensitivity.532REIMPLEMENT
Engage with Chapter leads and business stakeholders to determine priority Reference Architecture / blueprints and create a backlog (outside of Project Driven deliverables).522REIMPLEMENT
Uplift current Platform Architectural documentation to accurately reflect the implemented platform.542REIMPLEMENT
Create an engagement process where Solution Architecture can be endorsed and centralised for new projects to reduce rework.533REIMPLEMENT
Create solutions module repository of common reference Terraform modules for deploying Azure resources.532REIMPLEMENT

8.4          Azure DevOps

8.4.1      Definition

Azure DevOps is Microsoft’s SaaS offering for a holistic DevOps toolchain. It provides developer services to support teams to plan work, collaborate on code development and build/deploy applications and infrastructure. Azure DevOps is comprised of 5 core service functions:

Service NameDescription
Azure ReposProvides Git repositories or Team Foundation Version Control (TFVC) for source control of your code
Azure PipelinesProvides build and release services to support continuous integration and delivery of your apps
Azure BoardsDelivers a suite of Agile tools to support planning and tracking work, code defects, and issues using Kanban and Scrum methods
Azure Test PlansProvides several tools to test your apps, including manual/exploratory testing and continuous testing
Azure ArtifactsAllows teams to share Maven, npm, and NuGet packages from public and private sources and integrate package sharing into your CI/CD pipelines

8.4.2      Current State

CUST have implemented Azure DevOps with the One Organisation, Many Projects, Many Teams principle, which is the implementation model recommended, this both reduces significant operational overhead from a platform teams’ perspective whilst balancing the most enablement for application/development teams.

The DevOps team will be responsible for Landing Zone build which means each onboarding appplication/team will have their own Project, Service connection and Managed identity to deploy infrastructure, which creates a clear delineation between platform and application teams.

Platform repo’s make use of feature and fix branches for both complex changes and there is a branch policy on the ‘Master’ branch which requires 1 approval to merge, however, the policy does allow requestors to approve their own changes. The policy also has comment resolution checks enabled, however, comment resolution is optional meaning that pull requests could be commited with unresolved comments.

In the context of automation, the DevOps pipelines and Terraform code leverage a storage account as the backend for all Platform state.

Whilst this is good practice, there are a few configurations that require changes to have this particular storage account in a best-practice state:

  • Configure ABAC rules on a container level to only allow access to .tfstate files
  • Disable SAS keys and access tokens, enforcing Entra ID and RBAC only
  • Enable Blob versioning
  • Extend soft-delete to 30 days
  • Enable last accesss tracking time

8.4.3      Cloud Adoption Framework (CAF)

The CAF has a dedicated section for the use of Platform automation and DevOps, covering how most traditional IT operating models are not compatible with the cloud. Requiring an operational and organisational transformation to delivery against what are most likely significant enterprise migration targets. It is there for recommended to use a DevOps based approach for both the delivery of applications and for that of a central IT capability.

Landing Zones are a concept introduced via the Cloud Adoption Framework, defined as, the output of a multi-subscription Azure environment that accounts for scale, security, governance, networking, and identity boundaries. Essentially creating a repeatable deployment space for applications and infrastructure meeting pre-defined enterprise standards.

Outlined within the CAF are definitions of responsibilities for cross-functional DevOps teams such as Platform, Security, Networks and Application. In addition to this there is importantly a definition of central and federated responsibilities for the application and central platform teams. It is important to ensure that the application teams are empowered to control aspects of the application as part of a shared responsibility model.

8.4.4      Azure DevOps Recommendations

Whilst Azure DevOps is a product, it is important to note that the creation of a cloud-native operating model is required to take full advantage of any investment made towards an SRE or DevOps capability. Without such an operating model, teams are often hit obstacles related to lack of empowerment from senior leadership as there is no define cadence or rhythm.

Microsoft provide a design guide for the 3 main implementations of Azure DevOps:

  • One Organisation, One Project, Many Teams
  • One Organisation, Many Projects, Many Teams
  • Many Organisations, Many Projects, Many Teams

CUST have implemented the One Organisation, Many Projects, Many Teams, which is the implementation model I would recommend, this both reduces significant operational overhead from a platform teams’ perspective whilst balancing the most enablement for application / development teams.

Azure DevOps Pipelines can apply granular levels of RBAC controls based on a directory structure. To maximise the benefit of hierarchical inheritance this should be taken advantage of within scoping application deployments. This will further reduce the operational overhead of the central platform team managing Azure DevOps but also open opportunity for empowerment to scoped application or project teams.

Azure DevOps Recommendations

–       Define architecture of Landing Zones in Alignment with Reference Architectures.

–       Reconfigure the Master branch policy to require additional approvals and enforce comment resolution.

–       Reconfigure state storage account to be in line with the best-practice configuration.

RecommendationBenefitEffortComplexity

Reimplement/

Refactor

Define architecture of Landing Zones in Alignment with Reference Architectures.522REFACTOR
Reconfigure the Master branch policy to require additional approvals and enforce comment resolution.511REFACTOR
Reconfigure state storage account to be in line with best-practice configuration511REFACTOR

9            APPENDIX

9.1          Service Enablement

The Service Enablement Framework for Landing Zones on Microsoft Azure helps organisations in the regulated sectors meet compliance and security requirements while accelerating digital transformation. This framework provides a structured approach to defining, mapping, and enforcing necessary controls, balancing business needs with regulatory demands. It includes a prescriptive architecture and implementation plan, leveraging Azure services to streamline the process and ensure compliance with standards like PCI DSS, NIST 800-53, and SOC 1, 2, 3.

The framework outlines an operating model with clear separation of duties, involving Platform DevOps for operationalising the Azure platform and DevOps/AppOps for managing workloads within landing zones. Key functions include system operations, automation, and control mapping, all aimed at enabling broad adoption of Azure services while maintaining compliance and security. The steps involved are:

  • Define and map necessary controls for compliance and security.
  • Balance business needs with regulatory demands.
  • Implement a clear separation of duties between Platform DevOps and DevOps/AppOps teams.
  • Utilize Azure services to streamline the enablement process.
  • Ensure compliance with standards e.g.  PCI DSS, NIST 800-53, and SOC 1, 2, 3.
  • Follow a prescriptive architecture and implementation plan.
  • Map controls to Azure services and enforce them.
  • Differentiate between controls managed by Microsoft and those managed by platrform.
  • Map regulatory requirements to specific Azure services and controls.
  • Design the architecture to meet compliance and security needs.
  • Implement the designed architecture and controls.
  • Collect and maintain evidence of compliance.
  • Provide examples to illustrate the implementation process.
  • Outline the next steps for ongoing compliance and security management.

These steps ensure a comprehensive approach to service enablement, helping financial services organizations leverage Azure while maintaining stringent compliance and security standards.

To achieve this, the below images gives a high-level approach to initiating a request for an Azure service, Developing controls and managing ongoing requests using a defined process:

Service Enablement – FSI

Service Enablement Framework – Microsoft

  
  
  
  
  
  

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.