Solution Design
Azure Cloud Platform |
Contents
3.1.1 Compliance Requirements. 12
3.1.2 Functional Requirements. 12
3.1.3 Non-functional Requirements. 13
3.2 Architecture Principles. 14
4 Conceptual Solution Architecture. 20
4.1.1 Capability Model Reference Architecture. 22
4.1.2 Conceptual Application Architecture. 23
4.2 Target State – Solution Overview.. 24
4.2.1 Capability Model View.. 25
4.2.2 Standard Tooling Identification. 26
5.4 Platform Automation and DevOps. 30
5.5 Identity and Access Management. 31
5.5.2 Multi-factor Authentication. 31
5.5.3 Privilege Management. 32
5.5.5 Azure identity services tooling. 33
5.6 Public Key Infrastructure. 33
5.6.1 Internal Certification Authorities. 33
5.6.2 External Certification Authorities. 34
5.7.1 Networking Principles. 35
5.7.2 Azure network service tooling. 35
5.8.1 Azure secrets management tooling. 39
5.8.2 Data Encryption Keys. 40
5.9.1 Data-at-rest Encryption. 41
5.9.2 Data-in-transit Encryption. 41
5.9.3 Azure storage services tooling. 41
5.10.1 Azure compute service tooling. 43
5.11 Containers and Container Orchestration. 44
5.11.3 Azure container tooling. 47
5.12.1 Azure remote access tooling. 49
6.4 Azure resiliency tooling. 51
6.4.1 Recovery services vault. 51
7.1 Compliance Framework Alignment. 52
7.1.1 Cloud security posture management. 52
7.2 Security Log Aggregation. 53
7.3 Security Event and Incident Management. 53
7.4 Web Application Firewall 54
7.5 Cloud Workload Protection. 55
7.6 Cloud-native Application Protection. 55
7.7 Virtual Machine Endpoint Protection. 56
7.8 Vulnerability Management. 56
9 Monitoring Customer Experience. 60
Tables
Table 2 – RLX Contact Details. 2
Table 3 – RLX Version Control Details. 2
Table 4 – Compliance Requirements. 12
Table 5 – Functional Requirements. 13
Table 6 – Non-functional Requirements. 14
Table 7 – Architecture principles. 15
Table 8 – Solution Design Decisions. 19
Table 9 – Conceptual Design Principles. 21
Table 10 – Solution Standard Tooling. 26
Table 11 – Solution Mandated Tooling. 26
Table 12 – Networking Principles. 35
Table 13 – Azure Kubernetes Configuration Options. 45
Figures
Figure 1 – Azure Landing Zone conceptual diagram.. 20
Figure 2 – RLX Public Cloud Capability Model 22
Figure 3 – Application Conceptual Architecture. 23
Figure 4 – High-Level Architecture. 24
Figure 5 – Defined Capability Model Tool Mapping. 25
Figure 6 – Azure Hub-and-Spoke Topology. 34
Figure 7 – Key Vault capabilities. 39
Figure 8 – Contrast of Virtualisation Architectures Container Storage. 44
Design Decisions
Design Decision 1 – Subscription Sourcing Model 27
Design Decision 2 – Number of tenants in use. 28
Design Decision 3 – Structure of management groups. 28
Design Decision 4 – Use of subscriptions. 28
Design Decision 5 – Default language. 28
Design Decision 6 – Primary region for Identity. 29
Design Decision 7 – Primary Region for Production. 29
Design Decision 8 – Primary Region for Disaster Recovery. 29
Design Decision 9 – Primary Region for Nonproduction. 29
Design Decision 10 – Availability Zone Usage. 29
Design Decision 11 – Availability Zone Assignment. 29
Design Decision 12 – Infrastructure-as-Code. 30
Design Decision 13 – Platform Automation Tooling. 30
Design Decision 14 – CI/CD Agents. 30
Design Decision 15 – Version Identification. 30
Design Decision 16 – Identity Source. 31
Design Decision 17 – Authentication. 31
Design Decision 18 – Multi-factor Authentication. 32
Design Decision 19 – Authorisation. 32
Design Decision 20 – Separation of Duties. 32
Design Decision 21 – Directory Services Tooling. 33
Design Decision 22 – Certification Authorities (Internal). 34
Design Decision 23 – Certification Authorities (External). 34
Design Decision 24 – Network Topology. 37
Design Decision 25 – Route Management. 38
Design Decision 26 – Firewall Tooling. 38
Design Decision 27 – Firewall SKU.. 38
Design Decision 28 – HTTPS Inspection. 38
Design Decision 29 – Secrets Tooling. 40
Design Decision 30 – Volume Encryption Keys. 40
Design Decision 31 – Default Storage Replication. 40
Design Decision 32 – Default Disk Performance Tier. 40
Design Decision 33 – Storage Encryption Keys. 41
Design Decision 34 – Linux (RPM) Server Operating System.. 42
Design Decision 35 – Linux (Debian) Server Operating System.. 42
Design Decision 36 – Windows Server Operating System.. 43
Design Decision 37 – Windows Client Operating System.. 43
Design Decision 38 – Virtual Machine Image Management. 43
Design Decision 39 – Container Registries. 45
Design Decision 40 – Kubernetes Kernel Version. 45
Design Decision 41 – Kubernetes Network Stack. 45
Design Decision 42 – Kubernetes Network Controls. 46
Design Decision 43 – Kubernetes Deployment Architecture. 46
Design Decision 44 – Kubernetes Access. 46
Design Decision 45 – Container Deployment Configuration Management. 47
Design Decision 46 – Remote Access Tooling. 48
Design Decision 47 – Restricted Environment (CDE) Access. 48
Design Decision 48 – Privileged Access Environment. 48
Design Decision 49 – Availability Targets. 50
Design Decision 50 – Availability Architecture. 50
Design Decision 51 – Backup Tooling. 51
Design Decision 52 – Compliance Framework(s) Alignment. 52
Design Decision 53 – Cloud Security Posture Management Tooling. 53
Design Decision 54 – Security Log Aggregation. 53
Design Decision 55 – Security Incident and Event Management Tooling. 54
Design Decision 56 – Web Application Firewall (Internet) Tooling. 54
Design Decision 57 – Web Application Firewall (Internal) Tooling. 55
Design Decision 58 – Cloud Workload Protection Tooling. 55
Design Decision 59 – Cloud-native Application Protection Tooling. 55
Design Decision 60 – Virtual Machine Endpoint Protection Tooling. 56
Design Decision 61 – Monitoring Tooling. 58
Design Decision 62 – Operational Log Aggregation. 58
Design Decision 63 – Update management for Azure Virtual Desktop Session Hosts. 59
Design Decision 64 – Update management for Virtual Machines. 59
Design Decision 65 – Update management for Container images. 59
Design Decision 66 – Update management for Azure Kubernetes Services. 59
1 Executive Summary
This document details the solution approach to building, deploying, and operating a modern, secure Microsoft Azure platform for applicable business services, including internal and external facing functions and services.
Leveraging RLX’s experience across multiple public cloud platforms, Microsoft’s Cloud Adoption and Well-Architected Frameworks, national and regional governments, and secure-by-design core values, this document details the architecture framework and design requirements for this Azure tenant that will be used to:
- Enable highly scalable cloud-native workloads to be deployed and managed.
- Provide a secure platform for a mission-critical public service application.
- Provide assurance on intent and direction for cloud-native architecture.
There are several key outcomes of this and related documents, namely:
- Establishment of the Azure hierarchy of management groups, subscriptions, and resource groups.
- A highly available, scalable network to serve current and future demand.
- An integrated identity and access management solution, tied to a ‘least-privilege’ approach and supporting both first-party and third-party platform consumers.
- Clearly defined security and operational guardrails on the use of the platform.
- Adherence, visibility, and governance of applicable regulatory and compliance-aligned frameworks.
- Holistic reporting and monitoring across the platform.
- DevSecOps and platform automation at an enterprise level.
- Definition of capabilities to support disaster recovery and high availability requirements.
Business benefits delivered through the adoption of this architecture and related initiatives are:
- Well-known, near-real-time security posture from initiation through to operations.
- Defined performance and availability benchmarks.
- Rationalisation and/or lower complexity in
- Reduction in services and/or functions
- Innovation or agility in technology adoption and consumption
- Scalability, responsiveness, and flexibility
2 Introduction
2.1 Purpose
The purpose of this High-Level Design (HLD) is the creation of a Microsoft Azure Enterprise-Scale Landing Zone (ESLZ). The HLD describes the required Azure components and related services to operate an Azure ESLZ, as well as providing an overview of the specific tailoring for CUST’s business, functional, and non-functional technical requirements.
The overall architecture for the Azure environment, including this Landing Zone, is intended to follow a pattern-based approach that looks to apply enterprise-wide principles to Azure platform design and implementation, together with best practices from the Azure Cloud Adoption and Well-Architected Frameworks.
The architecture patterns applied to this Landing Zone have been selected with the following outcomes:
- A secure, sustainable digital infrastructure using Infrastructure-as-Code (IaC), immutable infrastructure, and separation of duties through DevSecOps.
- A modern, future-forward seamless digital platform that facilities innovation using a “SaaS over PaaS over IaaS” philosophy for workloads and elasticity.
- Providing a solid technology-based foundational platform for future growth into Azure for CUST ’s internally and externally facing business applications.
- Efficient, cost-effective, and scalable tooling, leveraging secure-by-design principles and cloud-native technologies.
- Ensuring continuity of service for all platform consumers by embedding the availability and resilience of the Azure platform.
- Ensuring continuity of projects and development, using highly available Software- and Platform-as-a-Service offerings, that are consumption-based and benefit from out-of-the-box security and resiliency.
The HLD also provides a common understanding of the decisions, their rationale, and tooling used to operate the ESLZ and related services to all stakeholders.
2.2 Scope
The components listed below are in scope of this HLD for the Azure ESLZ:
- Enterprise Agreement enrolment and Entra ID tenant(s)
- Entra ID Directory-based identity and access management
- Azure management groups and subscription hierarchies
- Network topology, integration, and delegation
- Reporting, monitoring, and operational capabilities
- Platform-wide security configuration, governance, and compliance
- Platform-wide automation and DevSecOps capabilities
2.3 Out of Scope
Any item not identified in Section 2.2 will be considered out of scope for the purpose of this document.
3 Design Approach
3.1 Critical drivers
3.2 Solution principles
ID | Description |
SP-1 | Secure environment for research |
SP-2 | Alignment with Australian Government ISM to PROTECTED |
SP-3 | Rapid spin-up of services |
SP-4 | Ephemeral workloads |
SP-5 | Scalable and cost effective – template driven |
SP-6 | High core compute storage capacity and processing power |
SP-7 | Consistent and modular policy and governance – greenfield policies and risk management practices |
3.3 Requirements
3.3.1 Compliance Requirements
The data and workloads hosted through the platform are required to comply with several security and regulatory frameworks. These frameworks, and high-level alignment requirements are listed below.
ID | Requirement | Description |
CR-1 | Platform-wide classification | The platform, including the ESLZ, will be built to known-good patterns, which have been designed to meet a stringent controls framework such as the ISM. |
CR-2 | Data sovereignty | The platform, including the ESLZ, is required to be hosted within Australian borders. |
CR-3 | Regulatory compliance | The platform, including the ESLZ, is required to meet the following regulatory framework(s): – IRAP PROTECTED |
CR-4 | Security compliance | The platform, including the ESLZ, is required to align to the following security compliance framework(s): – ACSC Essential 8 – IRAP PROTECTED – Secure Cloud Computing Architecture (SCCA) |
Table 4 – Compliance Requirements
3.3.2 Functional Requirements
This section outlines the functional requirements identified for the Azure platform. Functional requirements are the capabilities related to the performance or operation required of a given system, solution, or capability.
ID | Requirement | Description |
FR-1 | Scalable hosting | The platform should be capable of hosting any number of services desired to be served through the ESLZ. |
FR-2 | Enterprise monitoring | The platform should provide an end-to-end monitoring system for every service of the ESLZ and other workloads. |
FR-3 | Resiliency | The platform should provide end-to-end recovery and resilience tooling to ensure applications are able to meet agreed availability and recovery service levels. |
FR-4 | Reporting | The platform should provide the ability to report on workloads, performance, cost, security, and configuration. |
FR-5 | Administration | The platform should provide all administrative capabilities through a central interface or portal. |
FR-6 | Language | The platform is required to support the English language as a default, with other non-English languages being optional for support reasons. |
FR-7 | Support | The platform must support the ability for wide range of administrative or privileged users to be able to connect to the environment to support its’ operation. |
Table 5 – Functional Requirements’
3.3.3 Non-functional Requirements
This section outlines the non-functional requirements identified for the Azure platform. Non-functional requirements are the inherent characteristics unrelated to perform of a given system, solution, or capability.
ID | Requirement | Description |
NR-1 | Cost efficiency | The platform should minimise cost incurred to support the platform, by adopting: – Dynamic scaling based on demand. – Minimise or eliminate idle resources. – Deploy small and scale up/out. |
NR-2 | Standardisation | The platform should deploy resources based on existing technology standards as well as from approved solution architecture patterns. |
NR-3 | Operationalisation | The platform should support enacting change at scale without impacting existing operations. |
NR-4 | Compliance | The platform should support dynamic compliance based on workload requirements, as well as platform-wide requirements. |
NR-5 | Availability | The platform should provide patterns and capabilities to support achieving the following targets: – The ESLZ must achieve availability of 99.9% at go-live. |
NR-6 | Scalability | The platform should provide patterns and capabilities to enable the following approaches for scaling workloads: – Vertical scaling – Horizontal scaling – Automatic scaling |
NR-7 | Security | The platform should provide patterns and capabilities which include: – Least-privilege permission(s) – Most-secure configuration(s) |
NR-8 | Simplicity | The platform should provide patterns and capabilities which follow “least-complexity” |
NR-9 | Identification of users | The platform needs to identify unique users who are consuming the platform for communication, administrative and/or privileged functions. |
Table 6 – Non-functional Requirements’
3.4 Architectural approach
3.4.1 Microsoft frameworks
Microsoft offers several frameworks to support build and operations of cloud environments. These are:
- Cloud Adoption Framework (CAF)
- Well-Architected Framework (WAF)
The CAF is specifically designed to support new organisations and enterprises prepare for cloud-based workloads and ways of working. The WAF is designed to support the refactoring or build of application on cloud, to best maximise the value of cloud. Both of these frameworks are leveraged in the build of the Enterprise-Scale Landing Zone.
The primary purpose of the Azure Enterprise-Scale Landing Zone (ESLZ) is to ensure that when a new workload is deployed to Azure, such as a new secure research project, the supporting foundation is already in place which reduces onboarding time and provides compliance in line with organisational security and governance requirements.
RLX will create a new Azure ESLZ in an existing Microsoft Entra ID tenancy. This involves prescriptive guidance coupled with best practice for the Azure control plane. Azure Landing Zones are predefined and preconfigured environments which provide baselines for scalability, security, governance, networking, and identity. Azure Landing Zones will be built to align with CUST’s operational requirements.
The Azure ESLZs to be provisioned contains the following architectural components:
- Scalable and cost-optimised hub-and-spoke network topology using logically segmented subnets, on-premises connectivity and controlled routing behaviours
- Pre-defined governance, monitoring, compliance, and security controls
- DevOps platform to support automated infrastructure deployments and maintenance updates
- Integration with enterprise identity platform to control environment access and manage user lifecycle
- Centralised billing with the ability to support chargeback scenarios
3.4.2 Architecture principles
The following principles have been applied to this Design and related Designs:
Principle | Description |
Business focussed | The solution is aligned to business requirements and scenarios over technical requirements. |
Future-forward | The solution will maximise default features of the platform and use standardised industry practices for configuration or automation activities to support the advancements/evolution of the platform. |
Implementation agility | Tactical solutions are used when articulated strategic solutions are not ready when design commences. Microsoft native first-party software may be used as a fall back when the strategic solution(s) are not ready. |
Native tooling | Tools native to the platform is desirable over third-party solutions; balanced against tooling aligned to technology or security strategies. |
Secure-by-design | All components, services, or capabilities must implement security controls designed to improve the confidentiality, availability, and integrity of the workloads residing in the platform, and the platform itself. |
Authenticated | All consumers of the platform and hosted workloads must be identifiable back to a given, unique, named user identity. |
Authorised | All consumers of the platform and hosted workloads must have the minimum possible privileges associated to their identity, with the option to upgrade their privileges through an approved automated process. |
Resilient | The platform can gracefully handle and recover from failures during or following event types: disruptions, failures, and incidents. |
Available | The platform is built with redundancy to avoid having single points of failure. |
Performant | The platform is built in a way that always ensures consistent performance of key components, even under extended intensive use. |
Scalable | The platform can support dynamic utilisation and availability, regardless of the number of workloads or environments operating. |
Demand-driven capacity | The platform is designed to operate at a minimal capacity level by default and scale up & out to support higher demand rather than running at a fixed capacity. |
Operable | The platform maximises Microsoft – where reasonable – and supports third-party tooling (where possible) to enable management teams to support the platform |
PaaS-first | The platform will favour Platform-as-a-Service (PaaS) over Infrastructure-as-a-Service (IaaS) services where possible. |
Data tiering | The platform will leverage the most appropriate data storage tier for the workload based on performance, manageability, security, and cost requirements. |
Continuous integration | The platform maximises automation to reduce deployment times and minimise human/manual prone errors. Infrastructure is treated as code and released through automation, where errors or misconfigurations are caught as part of development, not after deployment. |
Table 7 – Architecture principles
3.5 Assumptions
3.6 Design Decisions
The below table captures the results of each Design Decision. For the detail in each Decision, refer to the specific Decision in the design.
Decision # | Title | Decision | Section |
DD-ELZ-1 | Subscription Sourcing Model | An Enterprise Agreement will be used to facilitate subscription billing. | 5.1 |
DD-ELZ-2 | Number of tenants in use | Only a single tenant will be deployed. | 5.2 |
DD-ELZ-3 | Structure of management groups | Management groups will be used to group subscriptions together logically. | 5.2 |
DD-ELZ-4 | Use of subscriptions | Subscriptions will be used to segregate resources per functional domain and environment. | 5.2 |
DD-ELZ-5 | Default Language | The default language of the tenant will be set to ‘English’. | 5.2 |
DD-ELZ-6 | Primary Region for Identity | Australia East will be the location used for Entra ID | 5.3 |
DD-ELZ-7 | Primary Region for Production | Australia Central will be the Azure region used for hosting production environments. | 5.3 |
DD-ELZ-8 | Primary Region for Disaster Recovery | Australia Central 2 will be the Azure region used for hosting disaster recovery environments. | 5.3 |
DD-ELZ-9 | Primary Region for Nonproduction | Australia Central will be the Azure region used for hosting nonproduction environments. | 5.3 |
DD-ELZ-10 | Availability Zones Usage | Not applicable for the environment given DD-ELZ-7 and DD-ELZ-8 | 5.3 |
DD-ELZ-11 | Availability Zones Assignment | Not applicable for the environment given DD-ELZ-7 and DD-ELZ-8 | N/A |
DD-ELZ-12 | Infrastructure-as-Code | Multiple infrastructure-as-code languages will be in use | 5.4 |
DD-ELZ-13 | Tooling | A new Azure DevOps environment will be used as the version control system for Solution Design-related actions, such as infrastructure-as-code and CI/CD. | 5.4 |
DD-ELZ-14 | CI/CD agents | Local CI/CD agents will be deployed into the environment for secure deployments | 5.4 |
DD-ELZ-15 | Version Identification | Semantic Versioning will be used to identify versions of code in use, for infrastructure and configuration. | 5.4.1 |
DD-ELZ-16 | Identity source | Entra ID will be the source of truth for all identities. | 5.5.1 |
DD-ELZ-17 | Authentication | Entra ID will be the source of truth for all authentication consumers. | 5.5.1 |
DD-ELZ-18 | Multi-factor authentication | Multi-factor authentication using Entra ID will be required on all accounts, both administrative and otherwise. | 5.5.2 |
DD-ELZ-19 | Authorisation | Entra ID will be the source of truth for all authorisation activities. | 5.5.3 |
DD-ELZ-20 | Separation of Duties | Separate privileged and non-privileged accounts will be used to isolate. | 5.5.3 |
DD-ELZ-21 | Directory Services Tooling | Entra ID Domain Services will be deployed to the environment. | 5.5.4 |
DD-ELZ-22 | Certification Authorities (Internal) | An internal PKI environment will be established within the Azure environment. | 5.6.1 |
DD-ELZ-23 | Certification Authorities (External) | The preferred external, publicly trusted, certificate issuer is Let’s Encrypt. | 5.6.2 |
DD-ELZ-24 | Topology | A hub-and-spoke network topology will be adopted and implemented. | 5.7.1 |
DD-ELZ-25 | Route Management | User-defined Routes (UDR) will be used to control network routing | 5.7.1 |
DD-ELZ-26 | Firewall Tooling | Azure Firewall will be leveraged as the network policy and routing management tool. | 5.7.1 |
DD-ELZ-27 | Firewall SKU | Azure Firewall Premium will be the SKU selected for Azure Firewall | 5.7.1 |
DD-ELZ-28 | HTTPS Inspection | Outbound TLS inspection will be enabled. | 5.7.1 |
DD-ELZ-29 | Secrets Tooling | Adopt Azure Key Vault for Solution Design secret, key, and certificate management. | 5.8 |
DD-ELZ-30 | Volume Encryption Keys | Implement Customer-managed keys stored in Key Vault for volume encryption. | 5.8.1 |
DD-ELZ-31 | Default Storage Replication | Locally-Redundant Storage (LRS) will be the default replication tier for storage volumes | 5.9 |
DD-ELZ-32 | Default Disk Performance Tier | Standard SSD will be the default performance tier for virtual machine disks. | 5.9 |
DD-ELZ-33 | Storage Encryption Keys | Adopt and implement custom keys for storage and data encryption | 5.9.1 |
DD-ELZ-34 | Linux (RPM) Server Operating System | IaaS VM Operating System version specification as Red Hat Enterprise Linux 9, or newer. | 5.10 |
DD-ELZ-35 | Linux (Debian) Server Operating System | IaaS VM Operating System version specification as Canonical Ubuntu Server 22.04, or newer LTS iterations. | 5.10 |
DD-ELZ-36 | Windows Server Operating System | IaaS VM Operating System version specification as Windows Server 2022, or newer. | 5.10 |
DD-ELZ-37 | Windows Client Operating System | IaaS VM Operating System version specification as Windows 11 Professional, or newer. | 5.10 |
DD-ELZ-38 | Virtual Machine Image Management | Hashicorp Packer will be used to develop, create, and publish Virtual Machine images. | 5.10 |
DD-ELZ-39 | Container Registries | One container registry instance will be deployed per environment. | 5.11.1 |
DD-ELZ-40 | Kubernetes Kernel Version | The minimum version of Kubernetes that will be deployed to the environment will be 1.27.1. | 5.11.2 |
DD-ELZ-41 | Kubernetes Network Stack | Kubernetes clusters will deploy the ‘Azure CNI’ network stack. | 5.11.2.1 |
DD-ELZ-42 | Kubernetes Network Controls | Kubernetes clusters will implement Calico network policies. | 5.11.2.2 |
DD-ELZ-43 | Kubernetes Deployment Architecture | Kubernetes clusters will deploy as ‘private’ clusters. | 5.11.2.3 |
DD-ELZ-44 | Kubernetes Access | Kubernetes clusters will implement RBAC and restrict access to specific namespaces. | 5.11.2.4 |
DD-ELZ-45 | Container Deployment Configuration Management | Helm will be used to manage deployment configuration(s) for containers on Kubernetes | 5.11.3 |
DD-ELZ-46 | Tooling | Remote access into the environment will be provided through an Azure Virtual Desktop deployment. | 5.12 |
DD-ELZ-47 | Restricted Environment Access | Privileged access within the environment will be through a separate Azure Virtual Desktop deployment. | 5.12 |
DD-ELZ-48 | Privileged Access Environment | Privileged access within the environment will be through a separate Azure Virtual Desktop deployment. | 5.12 |
DD-ELZ-49 | Availability Targets | Essential Solution infrastructure (IaaS) services will have an availability target of 99.95% (3-and-a-half nines). | 6.1 |
DD-ELZ-50 | Availability Architecture | Availability architectures will be built to support active/passive capability by default. | 6.2 |
DD-ELZ-51 | Backup Tooling | Azure Recovery Services Vault will be the default backup and replication tool for Azure-hosted resources. | 6.3 |
DD-ELZ-52 | Compliance Framework(s) Alignment | Configure resources and services to comply with one or more compliance frameworks. | 7.1 |
DD-ELZ-53 | Cloud Security Posture Management Tooling | Lacework will be deployed as the CSPM solution for the environment. | 7.1.1 |
DD-ELZ-54 | Security Log Aggregation | Azure Log Analytics will be used to centrally capture and store security event logs, such as Entra ID audit logs or Azure activity logs. | 7.2 |
DD-ELZ-55 | Security Incident and Event Management Tooling | The existing Microsoft Sentinel instance will be re-used as the SIEM solution for the environment. | 7.3 |
DD-ELZ-56 | Web Application Firewall Tooling | WAF services have not been identified as a requirement for this environment. | 7.4 |
DD-ELZ-57 | Web Application Firewall Tooling | WAF services have not been identified as a requirement for this environment. | 7.4 |
DD-ELZ-58 | Cloud Workload Protection Tooling | Lacework will be deployed as the CWPP solution for the environment. | 7.5 |
DD-ELZ-59 | Cloud-Native Application Protection Tooling | Lacework will be deployed as the CNAPP solution for Kubernetes and containers. | 7.6 |
DD-ELZ-60 | Virtual Machine Endpoint Protection Tooling | Defender for Endpoint will be deployed as the EDR and NGAV solution for virtual machines. | 7.7 |
DD-ELZ-61 | Vulnerability Management | Vulnerability assessment of containers will leverage Lacework. | 7.8 |
DD-ELZ-62 | Vulnerability Management | Vulnerability assessment of IaaS VMs will leverage Lacework. | 7.8 |
DD-ELZ-63 | Monitoring Tooling | Azure Monitor will be used as the default monitoring solution for the environment. | 8.1 |
DD-ELZ-64 | Operational Log Aggregation | Azure Log Analytics will be used to centrally capture and store operational logs and metrics, such as Container Insights, for Solution Design management services. | 8.2 |
DD-ELZ-65 | Update management for Azure Virtual Desktop Session Hosts | Azure Virtual Desktop Session Hosts will be updated through a release management process | 8.2 |
DD-ELZ-66 | Update management for Virtual Machines | Azure Update Management will be used to manage updates on Virtual Machines | 8.2 |
DD-ELZ-67 | Update management for Container images | Container images will need to be updated in their build configuration | 8.2 |
DD-ELZ-68 | Update management for Azure Kubernetes Services | Azure Kubernetes Services will set to update automatically within the cluster configuration | 8.2 |
Table 8 – Solution Design Decisions
4 Solution Overview
The Azure Platform is the foundational infrastructure capability provider, from which one can design, build, deploy, and manage the full spectrum of applications required to support all business technology initiatives.
The below image shows the proposed concepts being the solution. Each subscription is focused on a specific capability; shared – such as identity – or dedicated. This meets the Enterprise-Scale Landing Zone (ESLZ) recommendation.
The Enterprise-Scale architecture provides prescriptive architecture guidance coupled with Azure best practices and follows design principles across the critical design areas for an organisation’s Azure environment and landing zones. It is an architecture approach and reference implementation that enables an effective adoption and implementation of landing zones on Azure that are easily built and scalable.
Figure 1 – Azure Landing Zone conceptual diagram
The Conceptual Architecture is closely aligned with the Microsoft Cloud Adoption Framework’s approach to an Enterprise-Scale Landing Zone and focuses on the following five design principals.
Design Principle | Impact(s) |
Subscription democratisation | – This principle suggests production operations transitioned to the business units and workload teams. – This allows workload owners to have more control and autonomy of their workloads within the guardrails established by platform foundation. |
Policy-driven governance | – By not utilising Azure Policies to create guardrails within the environment, it increases the operation and management overhead of maintaining compliance. – Azure Policies helps to restrict and automate the desired compliance state within your environment. |
Single control plane | – Choosing a multi-vendor approach to operate control and management planes might introduce complexity of integration and feature support. – Replacing individual components to achieve “best of breed” or multi-vendor operations tooling might have limitations and could cause unintended errors due to inherent dependencies. – For Service Providers who are bringing an existing tooling investment to operations, security, or governance, a review of the Azure services and any dependencies is recommended. |
Application-centric and archetype-neutral | – By segmenting the workloads following a structure that differs from the example shown in the conceptual architecture in the management group hierarchy (such as an organizational hierarchy structure or grouping by Azure service), one can create a complex policy and access control structure to govern the entire environment. – This trade off introduces the risk of unintentional policy duplication and thereby exceptions, which adds to operational and management overheads. |
Azure-native design | – Like “Single control and management plane”, by introducing third-party solutions into the Azure environment, a dependency is created upon the solution to provide feature support and integration with Azure first party services. |
Table 9 – Conceptual Design Principles
4.1 Capability Model Reference Architecture
The below diagram demonstrates RLX’s reference architecture and capability model approach for cloud architecture. As part of the design process, each desired capability is mapped to one or more technologies. Existing technologies which are leveraged for an environment are captured, reviewed, and determined if appropriate for use or if needing to be replaced with a tool which meets functional and non-functional requirements.
Figure 2 – RLX Public Cloud Capability Model
4.2 Conceptual Application Architecture
The following diagram illustrates how the Conceptual Architecture can be implemented to support applications at scale. Where the Identity, Connectivity, and Management components are “shared” across one or more workloads.
This enables each individual components can scale individually, based on pre-determined or dynamic conditions and with different mechanisms. In this way, cost/value benefit can be obtained as some shared services have a baseline capacity supporting a given number of services at any one time.
Figure 3 – Application Conceptual Architecture
4.3 Target State – Solution Overview
The target state hosting platform is the Azure Cloud. Azure IaaS and PaaS services will be used to provide the platform level support services. The Azure platform will be built from several key technologies and services, as described further within this Solution Design.
Figure 4 – High-Level Architecture
4.3.1 Capability Model View
The below Reference Architecture Capability Model identifies each capability expected to be consumed within the Azure platform.
Figure 5 – Defined Capability Model Tool Mapping
4.3.2 Standard Tooling Identification
The below table identifies the standard tooling set applied within the Azure platform, enabling key ESLZ capabilities.
Category | Technology | Description |
Hosting | Microsoft Azure | Azure is a Cloud Computing provider offering a wide range of IaaS and PaaS services. |
Internet Connectivity | Microsoft Azure | Azure provides a secure, highly available private connectivity solution through ExpressRoute, supporting multi-gigabit throughput and a natively highly available peering. |
Common Services | Azure Monitor | Azure Monitor provides end-to-end visibility of logs, performance metrics, health checks and system alerts across the solution. It also drives system performance reporting. |
Azure Backup | Azure Backup provides backup and restore capabilities for file stores, VMs, databases in the solution. | |
Version Control | Azure DevOps | Azure DevOps provides an ecosystem for the end-to-end lifecycle management of infrastructure-as-code. |
Security Services | Azure Firewall | Azure Firewall provides network security capabilities for both Layer 4 and Layer 7 traffic and integrates with other Azure services for monitoring and reporting. |
Entra ID | Entra ID provides a central Identity and Access Management system for developers and administrators. It provides advanced security features such as risk-based authentication and privileged identity management. | |
Azure Key Vault | Key Vault provides HSM-backed storage for secrets, keys, and certificates. | |
Defender for Cloud | Defender for Cloud provides CSPM, CWPP, and CNAPP capabilities with a specialty for Azure environments. |
Table 10 – Solution Standard Tooling
4.3.2.1 Mandated Tooling
The following table describes the tooling which is required to support one or more application stacks within the environment.
Category | Technology | Description |
Security | Darktrace | Darktrace is the enterprise solution for detecting user behaviour and informational analytics. |
Security | Palo Alto Firewalls | Whilst Azure Firewall is serving East/West traffic in Azure (that is, Azure to Azure), Palo Alto firewalls will serve North/South traffic (that is, Azure to outside Azure). |
Table 11 – Solution Mandated Tooling
5 Solution Overview
The Microsoft Cloud Adoption Framework (CAF) for Azure is a series of guidance articles, patterns, and design approaches to help organisations to create and implement business and technology strategies for the cloud. It provides best practices, documentation, and tools. Cloud architects, IT professionals, and business decision makers use this information to achieve their cloud adoption goals.
The CAF’s Enterprise-Scale Landing Zone (ESLZ) architecture represents the strategic design path and target technical state for an organization’s Azure environment. It continues to evolve alongside the Azure platform. Azure landing zones are the output of a multi-subscription Azure environment that accounts for scale, security, governance, networking, and identity.
This Solution Design leverages the standard Microsoft CAF ESLZ architecture by extending the pattern, based on RLX’s experience and prioritisation of secure-by-design principles.
5.1 Billing Model
Microsoft consumption elements, namely Microsoft 365 licensing and Azure subscriptions, can be consumed through two (2) approaches:
- An Enterprise Agreement with Microsoft, or
- A “Cloud Services Provider” (CSP) engagement with a third-party.
An Enterprise Agreement is a direct negotiation engagement with Microsoft, where funds or a budget are allocated to commercial agreement which is drawn down or consumed over time. This does have a minimum time commitment of three (3) years but can consolidate costs for all Microsoft and Azure services.
A CSP is a third-party to both Microsoft and the end-consumer, that can provision services to the end-consumer without a formal budget allocation against an Enterprise Agreement. CSPs allow for end-consumers to manage cost truly flexibly, rather than be locked into a commercial agreement with Microsoft that may require change over time.
Design Decision | |
DD-ELZ-1 | Subscription Sourcing Model |
Decision | An Enterprise Agreement will be used to facilitate subscription billing. |
Justification | CUST already has an existing relationship with Microsoft and leverages the EA model to support a single interface point to consume Microsoft 365 and Azure services. |
Design Decision 1 – Subscription Sourcing Model
5.2 Tenant Hierarchy
A ‘tenant’ in terms of Microsoft is a directory for identities. Microsoft licenses and Azure subscriptions are assigned to a specific tenant. This serves as the central point for all billing elements specific to that tenant.
An organisation may operate one (1) or more tenants, depending on the applicable use case, such as an enterprise with multiple business names to isolate identities and costs for those businesses.
Design Decision | |
DD-ELZ-2 | Number of tenants in use |
Decision | Only a single tenant will be deployed. |
Justification | Whilst multiple tenants could separate Production from Nonproduction at the authentication step, given the relatively small size of the organisation, this may represent too much overhead and complexity to manage. Subscriptions can be used instead to separate those environments, along with the use of Privilege Management for temporary elevated Production access. |
Design Decision 2 – Number of tenants in use
Design Decision | |
DD-ELZ-3 | Structure of management groups |
Decision | Management groups will be used to group subscriptions together logically |
Justification | Management groups allow for the hierarchy of subscriptions and management groups. This means that access policies, Azure policies, and the like to be applied at different levels. For example, global ‘Reader’ access applied at the parent management group allows for all child objects to inherit that assignment. |
Design Decision 3 – Structure of management groups
Design Decision | |
DD-ELZ-4 | Use of subscriptions |
Decision | Subscriptions will be used to segregate resources per functional domain and environment. |
Justification | Using multiple subscriptions, aligned to business structure, allows for a separation of cost breakdown far more easily. Owners can be assigned view access to their subscriptions to review cost, and not risk confusing costs that are owned by other owners. |
Design Decision 4 – Use of subscriptions
Design Decision | |
DD-ELZ-5 | Default language |
Decision | The default language of the tenant will be set to ‘English’. |
Justification | The geographical region this environment is to be deployed to is Australia. The national language of Australia is English. |
Design Decision 5 – Default language
5.3 Azure Regions
The Azure Platform physical architecture will leverage Azure datacentres which are in the Australia East and Australia Southeast regions. Australia Southeast is located within the metropolitan Melbourne area, and Australia East is located within the Sydney metropolitan area.
Design Decision | |
DD-ELZ-6 | Primary Region for Identity |
Decision | Australia will be the location used for the Entra ID tenant |
Justification | Australia is the only region which supports running Entra ID within Australian borders. This does not preclude deploying Azure resources in other Australian regions. |
Design Decision 6 – Primary region for Identity
Design Decision | |
DD-ELZ-7 | Primary Region for Production |
Decision | Australia East will be the Azure region used for hosting production environments. |
Justification | Australia East represents the Azure region with the greatest capacity and availability of services; all of which are rated to the PROTECTED level. |
Design Decision 7 – Primary Region for Production
Design Decision | |
DD-ELZ-8 | Primary Region for Disaster Recovery |
Decision | Australia Southeast will be the Azure region used for hosting DR environments. |
Justification | Australia Southeast is the paired region for Australia East. |
Design Decision 8 – Primary Region for Disaster Recovery
Design Decision | |
DD-ELZ-9 | Primary Region for Nonproduction |
Decision | Australia East will be the Azure region used for hosting nonproduction environments. |
Justification | Australia East represents the Azure region with the greatest capacity and availability of services; all of which are rated to the PROTECTED level. |
Design Decision 9 – Primary Region for Nonproduction
Design Decision | |
DD-ELZ-10 | Availability Zone Usage |
Decision | Availability Zones will be used to deploy resources where supported. |
Justification | Australia East is the only region that supports Availability Zones; these should be adopted, for IaaS in specific zones or PaaS as “AZ” SKUs, to ensure that services deployed remain operational in the event of a single zone being unavailable. |
Design Decision 10 – Availability Zone Usage
Design Decision | |
DD-ELZ-11 | Availability Zone Assignment |
Decision | Zones 1 and 3 will be used to deploy Production resources, where a zone is required to be defined. |
Justification | Identification of specific zones for deployments allows for consistency of approach. For resources that are “multi-zonal”, without zone assignment, this isn’t relevant. |
Design Decision 11 – Availability Zone Assignment
5.4 Platform Automation and DevOps
Azure DevOps is currently recommended by Microsoft as one of the preferred Version Control Systems. Azure DevOps is a fully managed SaaS platform that orchestrates pipeline and agents.
GitHub Enterprise is an alternative solution as GitHub Enterprise is equivalent to GitHub’s public service but is designed for private use by large-scale enterprise and small-to-medium software development teams where they wish to host their repositories behind a corporate firewall.
Design Decision | |
DD-ELZ-12 | Infrastructure-as-Code |
Decision | Multiple infrastructure-as-code languages will be in use |
Justification | Each IaC language has different capabilities and integrations. Bicep will be used to deploy Azure resources; Terraform will be used to manage Entra ID configuration; Packer will be used to build images; Helm will be used to package containers. |
Design Decision 12 – Infrastructure-as-Code
Design Decision | |
DD-ELZ-13 | Platform Automation Tooling |
Decision | A new Azure DevOps environment will be used as the version control system for Solution Design-related actions for infrastructure, such as infrastructure-as-code and CI/CD. |
Justification | To minimise the change management, given the short timelines, this existing tooling provides all the required functionality and capability to build the Azure environments. |
Design Decision 13 – Platform Automation Tooling
Design Decision | |
DD-ELZ-14 | CI/CD agents |
Decision | Local CI/CD agents will be deployed into the environment for secure deployments |
Justification | Version control systems and/or CI/CD systems frequently require access to systems behind a firewall. Instead of granting access directly over the Internet, locally hosted agents that poll out to the VSC or CICD system can be deployed to an environment to support these activities. |
Design Decision 14 – CI/CD Agents
5.4.1 Version Control
Use of version control is valuable in identifying configuration and change in configuration over time. There are several modern methods for tracking this configuration., such as Semantic Versioning or Date-based Versioning.
Design Decision | |
DD-ELZ-15 | Version Identification |
Decision | Semantic Versioning will be used to identify versions of code in use, for infrastructure and configuration. |
Justification | Semantic Versioning (Major.Minor.Patch) provides an effective solution for identifying what “version” of a given configuration is presently in operation. |
Design Decision 15 – Version Identification
5.5 Identity and Access Management
5.5.1 Authentication
Authentication refers to the process of verifying the identity of a user or a system. The Azure platform solution will validate unique identifiers for users and systems through the Entra ID identity system. Entra ID will enable the use of a common set of policies, practices, and protocols to manage the identities and trust of users and systems.
Additional authentication controls include:
- Strict password policies – e.g., strength and prevention of re-use – will apply to Entra ID accounts.
- Requests for new accounts to follow a request-based process for auditability. Entra ID roles for account creation will be limited to an appropriate Entra ID group to prevent unauthorized creation of accounts.
- As a preventative measure against catastrophic administrator lockout from the Entra ID, emergency access accounts will be provisioned.
Design Decision | |
DD-ELZ-16 | Identity source |
Decision | Entra ID will be the source of truth for all identities. |
Justification | Entra ID is a cloud-first identity Solution Design, providing identity, authentication, authorisation, lifecycle, and privilege management capabilities. Entra ID can be extended to support LDAP integration points through Entra ID Domain Services. |
Design Decision 16 – Identity Source
Design Decision | |
DD-ELZ-17 | Authentication |
Decision | Entra ID will be the source of truth for all authentication consumers. |
Justification | Given Entra ID supports the use of single sign on (SSO), using SAML and OpenID Connect, privilege or role management will be centralised and simplified. Entra ID can be extended into the Azure environment through Entra ID Domain Services. |
Design Decision 17 – Authentication
5.5.2 Multi-factor Authentication
Entra ID will be configured to require at least two verification methods to confirm the identity of users. Policies will be in place to ensure that Entra ID based multi-factor authentication is enabled and enforced by default for all logins, including the Emergency Access Accounts (which should be stored in a physical secure location physically remote from the office location) and System Identities (used for system authentication).
Design Decision | |
DD-ELZ-18 | Multi-factor authentication |
Decision | Multi-factor authentication using Entra ID will be required on all accounts, both administrative and otherwise. |
Justification | Multi-factor authentication (something you know / something you have / somewhere you are) is a highly effective and efficient capability to prevent misuse of identities within an environment. |
Design Decision 18 – Multi-factor Authentication
5.5.3 Privilege Management
Entra ID provides identity governance capabilities through Entra ID Entitlement Management. Longitudinal access management is implemented via well-defined access packages and extending this pattern to delegate time-bound privileged access to resources, approvals, and access reviews, as well as facilitating access to external users where required.
Privileged Access Management (PAM) is required to control access to highly privileged credentials and prevent unauthorised access to the Azure platform. Entra ID Privileged Identity Management (PIM) will also be used to restrict access to privileged Azure resources. PIM provides users just-in-time privileged access to Azure and Entra ID resources, and oversight of what users are performing with their privileged access.
Design Decision | |
DD-ELZ-19 | Authorisation |
Decision | Entra ID will be the source of truth for all authorisation activities. |
Justification | Entra ID supports the ability to control levels of access granted to identities. This can be tied in with privilege management capabilities within Microsoft’s wider identity tooling portfolio, known as Microsoft Entra. |
Design Decision 19 – Authorisation
Design Decision | |
DD-ELZ-20 | Separation of Duties |
Decision | Separate privileged and non-privileged accounts will be used to isolate. |
Justification | Privileged accounts – except for break-glass accounts – should be restricted from being directly accessible from the Internet. Additionally, accounts used for communication such as emails, should be prevented from being able to access systems requiring privileged access. |
Design Decision 20 – Separation of Duties
5.5.4 Directory Services
Whilst Entra ID is a cloud-native solution, compatible with modern applications, there are several areas which Entra ID is not optimal as a solution. Entra ID Domain Services is an Azure service capable of extending Entra ID to fit those use cases, such as Kerberos authentication.
Design Decision | |
DD-ELZ-21 | Directory Services Tooling |
Decision | Entra ID Domain Services will be deployed to the environment. |
Justification | Based on the expected services required to support the operation of the environment, Entra ID Domain Services will be required to support central authentication and integration. |
Design Decision 21 – Directory Services Tooling
5.5.5 Azure identity services tooling
5.5.5.1 Entra ID
Entra ID is a cloud-first, cloud-native end-to-end identity and access management solution. It supports adoption and use of identities, conditional access policies governing how identities can be consumed, privilege management for controlled roles and/or group memberships. It also supports integration of applications that support Single-Sign On protocols, such as SAML v2, OpenID Connect, or OAuth 2.0.
5.5.5.2 Entra ID Domain Services
Entra ID Domain Services (AADDS) is an extension of Entra ID, to support traditional directory services for virtual machines and applications. An instance of AADDS provides the logical AD endpoints for connectivity, where the underlying infrastructure is entirely managed by Microsoft Azure.
5.6 Public Key Infrastructure
Public Key Infrastructure (PKI) is a term covering certificates, such as those used to enable HTTPS on websites, and the services that issue or sign certificates known as Certification Authorities (CAs). A known, defined PKI environment for internal and external certificate issuance is a key part of establishing trust on secured connections for data-in-transit security. Data-in-transit security is mostly performed through Transport Layer Security (TLS), which necessitates the adoption and implementation of trusted certificates.
5.6.1 Internal Certification Authorities
Internal CAs are required to be able to establish trust on an internal network, for services which require to expose or consume Transport Layer Security (TLS) protocols. This is an essential component for TLS inspection for firewall traffic.
Design Decision | |
DD-ELZ-22 | Certification Authorities (Internal) |
Decision | An internal PKI environment will be established within the Azure environment. |
Justification | The cost of using external trusted certificates for internal-facing applications can be ruinously expensive, both in terms of cost to procure and resource effort to manage. Self-signed certificates per resource would constantly show errors in web browsers and applications. An internal CA solution which allows for all systems to share a trusted issuer, along with automated enrolment eliminates those concerns to a major degree. This will be a requirement to support HTTPS inspection. The solution for PKI will be detailed within a PKI-specific design. |
Design Decision 22 – Certification Authorities (Internal)
5.6.2 External Certification Authorities
External CAs are required to be used to allow Internet-facing services to be trusted by end-users. The lifecycle management of issued certificates will be managed by a platform-hosted solution and integrated with a CA provider for zero- or low-cost certificates.
Design Decision | |
DD-ELZ-23 | Certification Authorities (External) |
Decision | The preferred external, publicly trusted, certificate issuer is Let’s Encrypt. |
Justification | Some external paid providers do allow for automated enrolment; however, the value in certificates is that it provides encryption for the communication channel. Most end-users (>95%) do not consider the provenance of a certificate as important, only that there is a certificate which does not throw errors. Let’s Encrypt supports automated enrolment with over the ACME protocol, with the certificate stored in Key Vault. |
Design Decision 23 – Certification Authorities (External)
5.7 Networking
Virtual Network(s) provide a key building block for establishing virtual private networks that isolate network communication within the Azure environment. By default, when an Azure resource is created and connected to Azure Virtual Network, they are allowed to route to any subnet within the virtual network, and outbound access to the Internet is provided by Azure’s Internet connection. This default configuration will be altered using User Defined Routes (UDRs) thus forcing Azure resource to route through virtual firewalls deployed in a central virtual network.
Figure 6 – Azure Hub-and-Spoke Topology
The network is designed with regional hubs and spoke based configuration to provide centralised management of core network concerns such as routing, firewalls, and DNS within each region.
Key design points include:
- Deployment across Australia East / Southeast paired regions for HA/DR.
- ExpressRoute private WAN connections to remote networks for performance and network security.
- Private network endpoints for Azure native services, to avoid sending sensitive internal traffic via the Internet.
The network aligns to a standard hub-and-spoke based architecture to enable centralised traffic management through the hub networks, and isolation of applications / workloads within the spoke networks. The Hub networks form the major connectivity transit zone for traffic within an environment. The current network architecture includes two hubs, one for production-based traffic and workloads as well as non-production network to enable less strict change control for development and testing purposes without impacting the production infrastructure.
5.7.1 Networking Principles
The following principles will be used to drive the selection of IP ranges and subnets for use within the environment.
ID | Principle | Rationale |
1 | Each region or environment must be allocated with an equivalently sized network. | This will allow for natural growth of the environment and synchronized network allocation for each region. |
2 | Each region must be allocated a minimum network capacity to support multiple capabilities. | This will support micro-segmentation and zero-trust principles, as well as initial growth-over-time. |
3 | Network allocations across regions and environments must not conflict. | This will ensure that if networks need to be integrated, no additional overhead needs to be applied for translation. |
4 | Network allocations across regions should mirror each other. | This will ensure that allocations are standard across environments. |
Table 12 – Networking Principles
5.7.2 Azure network service tooling
5.7.2.1 Azure virtual networks
Virtual Networks provide the Layer 3 capability to issue IPv4 and/or IPv6 addresses to Infrastructure-as-a-Service resources, such as Virtual Machines. Virtual Networks can host one or more IP CIDR ranges, allowing a significant number of IP resources to be consumed.
Subnets are the capability within virtual networks sub-divide and micro-segment networks, to ensure that only relevant services are hosted in the same subnet.
5.7.2.2 Azure firewall
Azure Firewall is a PaaS solution provided by Microsoft which performs Layer 4 and Layer 7 security policy management for network traffic, akin to traditional or Next-Generation Firewalls. Azure Firewall does require the use of Public IPs for management, to communicate with the Azure management API; these will be procured through Azure via an IP prefix assignment.
5.7.2.3 Public IP prefixes and addresses
Public IP addresses are unique addresses assigned to devices and resources connected to the internet. A public IP prefix is a range of consecutive public IP addresses allocated to a tenant. Instead of assigning individual IP addresses, a prefix is used to issue IP addresses due to the many advantages.
Using a prefix allows scalability and efficient allocation of IP addresses, a range can be assigned to a single block, reducing fragmentation, and optimizing address utilisation. Streamlined networking is also facilitated with the use of prefixes, the prefix can be referenced instead of individual IP addresses separately for greater network control and reduced change-over-time in configuration.
5.7.2.4 Route tables
By default, network traffic in Azure Virtual Networks is routed directly between subnets and IP addresses. To support Zero Trust principles and maintain security, custom routing logic is required to be introduced into each virtual network’s subnets.
Route tables will be used to statically set traffic paths as dynamic cross-region route management is not required.
5.7.2.5 DNS zones
Azure DNS zones act as DNS zone files, providing zone hosting with a wide range of DNS record capabilities and integration. Private DNS Zones are intended for use within a private network, not a public or Internet-facing network. Azure provides the capability to integrate with virtual networks, allowing for both resolution of DNS records as well as automatically registering resources within the virtual network to the DNS zone.
5.7.2.6 Network security groups
Network Security Groups (NSGs) act as layer 4 network access control lists, but not as powerful as a next-generation firewall inspecting layer 7 traffic. NSGs are the last line of network defence in Azure Virtual Networks, able to be applied to both subnets and network interfaces.
5.7.2.7 Application security groups
Application security groups (ASGs) are used in conjunction with network interfaces and NSGs. ASGs allow for groups of resources to be created, so that NSG rules only need to identify the ASG as the source or destination, removing the need for IP addresses to be entered. Network Interfaces can be attached to one or more ASGs, meaning an interface could be subject to multiple rules.
5.7.2.8 Load balancers
Load Balancers (LBs) in Azure come in several types. Layer 4 LBs for handling TCP and UDP and either as public or private; or, the Application Gateway for Layer 7 routing of HTTP traffic.
5.7.2.9 ExpressRoute circuits
ExpressRoute circuits are a highly available, low latency, high bandwidth wide-area network (WAN) connection option to be able to integrate Azure networks into an existing or private network. Consuming an Azure Virtual Network Gateway, along with a connection through a subset of network providers, ExpressRoute is a viable alternative to VPN tunnels or mesh networks where latency, privacy, and security are concerns.
5.7.2.10 Virtual and local network gateways
Virtual Network Gateways (VNGs) are a network routing appliance, able to operate either as a VPN server – terminating IPSec and SSL-VPN tunnels – or as an ExpressRoute circuit peer. VNGs use the BGP network protocol to be able to manage complexity and routing at scale across tunnels and circuits.
Local Network Gateways (LNGs) are used only in conjunction with VNGs when deploying IPSec tunnels, for static routing. If an environment at the other end of an IPSec tunnel can’t handle BGP for dynamic routing, an LNG needs to be deployed to instruct the VNG what networks are available on the other end of the tunnel and via which destination address.
Design Decision | |
DD-ELZ-24 | Network Topology |
Decision | A hub-and-spoke network topology will be adopted and implemented. |
Justification | A hub-and-spoke topology allows for the use of centralised network controls to manage network traffic and provide zero trust capabilities. |
Design Decision 24 – Network Topology
Design Decision | |
DD-ELZ-25 | Route Management |
Decision | User-defined Routes (UDR) will be used to control network routing |
Justification | UDRs are static routes applied to networks. UDRs are zero-cost options to control routing at scale in an environment where routing is isolated and not expected to require dynamic routing. |
Design Decision 25 – Route Management
Design Decision | |
DD-ELZ-26 | Firewall Tooling (East/West) |
Decision | Azure Firewall will be leveraged as the East/West firewall stack for network policy and routing management tool. |
Justification | Azure Firewall is a native service and has feature-parity with the key features required for the environments to be secure. |
Design Decision 26 – Firewall Tooling (East/West)
Design Decision | |
DD-ELZ-27 | Firewall Tooling (North/South) |
Decision | Palo Alto Prisma NGFW will be leveraged as the North/South firewall stack for network policy, routing management, and TLS inspection for traffic exiting Azure. |
Justification | Azure Firewall is a native service and has feature-parity with the key features required for the environments to be secure. |
Design Decision 27 – Firewall Tooling (North/South)
Design Decision | |
DD-ELZ-27 | Firewall SKU |
Decision | Azure Firewall Premium will be the SKU selected for Azure Firewall |
Justification | Premium provides additional security and scaling capabilities, such as TLS inspection, threat intelligence, and IDPS, which is not present in lower tiers. |
Design Decision 28 – Firewall SKU
Design Decision | |
DD-ELZ-28 | HTTPS Inspection |
Decision | Outbound TLS inspection will be enabled on North/South firewalls. |
Justification | Several systems are required to support this environment, which is subject to specific compliance requirements, where there is a need to manage and identify allowed external traffic. |
Design Decision 29 – HTTPS Inspection
5.8 Secrets Management
Azure Key Vault will be used to store encryption keys, secrets such as passphrases, and certificates used for HTTPS, as well as managing their lifecycle. Key Vault then is the source of truth for all certificates, secrets, and keys throughout the platform. Applications may deploy one or more Key Vaults as needed to secure keys and secrets.
Figure 7 – Key Vault capabilities
In above diagram, a VM is leveraging a certificate stored in Key Vault for providing HTTPS on a website. This allows for the certificate to be managed and updated without needing administrative access to the underlying VM. In addition, the web and VM logs are being pushed to a storage account file share, which is encrypted at the Azure management layer using a key stored in the same Key Vault.
This is an example of the capabilities that Key Vault can bring to help separate the elements which improve security within a given Azure cloud environment.
5.8.1 Azure secrets management tooling
Azure Key Vault supports keys, secrets, and certificates. Keys meaning RSA or EC keypairs; secrets meaning any value such as an API token; and certificates meaning PKCS12-formatted keypair and signed certificate bundles. Other services, such as Azure Storage, depend on Azure Key Vault for use in managed encryption, opposed to Microsoft-managed encryption.
An advantage of leveraging Key Vault is the ability to not just store secret values, such as keys, but also generate keys as well. With this, keys generated and stored in Key Vault cannot be exported in clear-text, reducing the ability of administrators to store highly sensitive credentials outside of controlled areas. This is as-designed for Key Vault to prevent key exposure. Instead, administrators may request a “key signing” operation instead, where it is required. For private keys that are required to be imported into Key Vault, this should instead be governed by an appropriate process.
Where possible, Key Vaults will be deployed with all defined access within the IaC definition; however, in the event of issues or troubleshooting, some access policies may need to be manually updated. These can then be ‘reset’ by re-running the deployment pipeline.
Design Decision | |
DD-ELZ-29 | Secrets Tooling |
Decision | Adopt Azure Key Vault for platform secret, key, and certificate management. |
Justification | Azure Key Vault is a native service offered in Azure and supports out-of-the-box integration to Entra ID for access policies or RBAC. Key Vault can support keys (EC, RSA), secrets, and certificates. |
Design Decision 30 – Secrets Tooling
5.8.2 Data Encryption Keys
Design Decision | |
DD-ELZ-30 | Volume Encryption Keys |
Decision | Adopt and implement custom keys for storage and data encryption |
Justification | Azure Key Vault is a native service offered in Azure and supports out-of-the-box integration to Entra ID for access policies or RBAC. Key Vault can support keys (EC, RSA), secrets, and certificates. |
Design Decision 31 – Volume Encryption Keys
5.9 Storage
The Azure platform will require the use of storage for storing and retaining data which supports the platform functionality, including virtual machine backups and logging. Specifically, VM disks, storage accounts and Log Analytics workspaces. To enable clear isolation and ownership boundaries, services will be provisioned and logically allocated to a specific application or service requirement. This approach also ensures scalability, availability and security can be configured based on the targeted needs of each service.
Design Decision | |
DD-ELZ-31 | Default Storage Replication |
Decision | Zone-Redundant Storage (ZRS) will be the default replication tier for storage volumes |
Justification | Zone-Redundant Storage (ZRS) is only available in Australia East, as the only region with Availability Zones. LRS or GRS volumes may be used for specific use cases. |
Design Decision 32 – Default Storage Replication
Design Decision | |
DD-ELZ-32 | Default Disk Performance Tier |
Decision | Standard SSD will be the default performance tier for virtual machine disks. |
Justification | Standard SSD is sufficient for most use cases for virtual machine storage. Premium SSD should be limited to only those use cases where it is a requirement, such as IaaS virtual machines running database software. It also represents a significant cost saving in the operation of a given environment. |
Design Decision 33 – Default Disk Performance Tier
5.9.1 Data-at-rest Encryption
Data at rest (DaR) encryption refers to the encryption of data as it persists in storage. DaR encryption is a critical security control as it addresses security risks related to direct physical access to storage media, as the underlying data is not recoverable and cannot be changed without the configured encryption key. This makes it an important layer in the defence in depth strategy of Microsoft data centres. In addition, there are often compliance and governance reasons to deploy DaR encryption.
There are two layers of data at rest security:
- Encryption managed by the Azure platform provider (Microsoft), and,
- Encryption managed by the Customer.
Out-of-the-box encryption as part of the Azure platform is managed by Microsoft. The storage account service encrypts data at-rest with the AES-256 encryption cipher.
All forms of storage will be encrypted at rest as well as in transit. The specific encryption protocols and resource configuration will be defined in the Detailed Design and enforced via Cloud Security Posture Management (CSPM) tooling.
Design Decision | |
DD-ELZ-33 | Storage Encryption Keys |
Decision | Customer-managed Keys will be used to encrypt storage volumes. |
Justification | Microsoft-managed keys are keys managed by Microsoft. For a more secure environment, keys created in the environment and stored in Key Vault can be leveraged to encrypt storage volumes instead. |
Design Decision 34 – Storage Encryption Keys
5.9.2 Data-in-transit Encryption
Data-in-transit (DiT) encryption refers to the encryption of data as it transits from one endpoint to another, usually across a network bearer. Broadly speaking DiT encryption is handled transparently by the Azure platform using TLS 1.2 or later both on external and internal Azure interfaces.
All the cloud services and protocols used for encryption are implemented should be using Australian Signals Directorate approved cryptographic algorithms. All PaaS services encrypt data in-transit with the most secure, supported Transport Layer Security protocol, either version 1.2 or 1.3.
5.9.3 Azure storage services tooling
5.9.3.1 Storage accounts
Storage accounts are a PaaS service offering, providing blob (block-level), file (SMB), table, and queue storage capabilities. Storage accounts can also be configured for replication tiers, with single region, multi-region, and/or multi-zonal support options.
5.9.3.2 Disks
Disks are an IaaS capability for mapping persistent storage to virtual machines to be treated as attached storage and is available in several tiers depending on performance requirements.
5.10 Compute
The compute aspects of Azure platform are based on two primary compute types, PaaS and Virtual Machines. The choice of technology is based on the availability of a suitable native technology, with a preference for PaaS over IaaS where possible.
Where possible, the preference is to utilise scale sets over specific Virtual Machine creation and management to enable additional scaling opportunities as well as to reduce the operation management and complexity of managing specific instances. To ensure consistency within Virtual Machine-based compute, Virtual Machines images will be built via automation tools, such as Hashicorp Packer.
The design decisions that follow provide a recommended approach on known, good, trusted Virtual Machine operating systems.
Design Decision | |
DD-ELZ-34 | Linux (RPM) Server Operating System |
Decision | IaaS VM Operating System version specification as Red Hat Enterprise Linux 9, or newer. |
Justification | Some applications require specific RPM kernel features, which are not present in Debian distributions, such as Oracle Databases. |
Design Decision 35 – Linux (RPM) Server Operating System
Design Decision | |
DD-ELZ-35 | Linux (Debian) Server Operating System |
Decision | IaaS VM Operating System version specification as Canonical Ubuntu Server 22.04, or newer LTS iterations. |
Justification | Ubuntu is a common Linux Solution Design that is used across several different use cases. Microsoft use Ubuntu as the Linux operating system underpinning Azure Kubernetes Services. |
Design Decision 36 – Linux (Debian) Server Operating System
Design Decision | |
DD-ELZ-36 | Windows Server Operating System |
Decision | IaaS VM Operating System version specification as Windows Server 2022, or newer. |
Justification | Windows Server 2022 has now been released for nearly two (2) years. It should now be the default for any new environments. |
Design Decision 37 – Windows Server Operating System
Design Decision | |
DD-ELZ-37 | Windows Client Operating System |
Decision | IaaS VM Operating System version specification as Windows 11 Professional, or newer. |
Justification | Windows 10 is end-of-support in 2025 with the latest release (22H2) being the final feature update the OS will receive. Security upgrades will continue until 2025 however most if not all applications will support running on Windows 11. |
Design Decision 38 – Windows Client Operating System
Design Decision | |
DD-ELZ-38 | Virtual Machine Image Management |
Decision | Hashicorp Packer will be used to develop, create, and publish Virtual Machine images. |
Justification | Packer is a standard open-source industry tool which is extremely effective at building virtual machine images. All configurations can be managed as code and updates to configuration managed through DevSecOps and CI/CD processes. This is the same process that Microsoft uses to create Azure DevOps agents. |
Design Decision 39 – Virtual Machine Image Management
5.10.1 Azure compute service tooling
5.10.1.1 Azure compute gallery
Azure Compute Galleries are the service which manages Standard Operating Environment images for virtual machines. Consumers create VMs using specific named images and versions. Image publishers create VMs, perform configuration, then publish the image to the gallery for re-use.
5.10.1.2 Azure virtual machines
The traditional IaaS capability to run a virtual machine, Azure supports both Windows and Linux VMs. There is an active marketplace of VMs, supporting vendor-provided images such as next-generation firewalls, or “complete” images such as a combined web/app/database all-in-one server.
5.10.1.3 Azure virtual machine scale-sets
VM Scale Sets are image-based VMs which scale-out and scale-in dynamically based on pre-defined conditions. Scale-sets can only run from an image and are not traditionally accessible like standard virtual machines. Scale-sets are best used for highly parallel workloads, such as data processing or web servers under heavy load.
5.10.1.4 Azure availability sets
Availability sets are a mechanism to separate virtual machines within Azure’s underlying physical infrastructure and management activities. When using more than 1 virtual machine, an availability set is used to identify fault domains, which VMs are then allocated into. In this way, application availability can be achieved, even when there is an issue with the underlying Azure infrastructure.
5.11 Containers and Container Orchestration
Containers are a modern standard for virtualisation of application hosting, which represents a second-generation approach to virtualisation. Under a well-known approach with Virtual Machines, applications were deployed to comparatively large instances, and exposed a “full” operating system for support. This allowed for a high degree of customisation and variability with identical applications, resulting in a varied support and performance baseline.
Contrasted to container-based applications, where an application is preconfigured with the same specification and performance benchmark. Rather than exposing a “full” operating system, a container ‘contains’ only the minimal binaries required to operate, enabling for a much greater volume of containers to be deployed.
Containerisation allows for the same application to be run in multiple different environments, agnostic to the underlying operating system of the host.
Figure 8 – Contrast of Virtualisation Architectures Container Storage[1]
Containers are like Virtual Machine images. They are compiled or built and need to be stored in preparation for deployment. Azure offers a native service, which supports containers.
Design Decision | |
DD-ELZ-39 | Container Registries |
Decision | One container registry instance will be deployed per environment. |
Justification | Separation of environments should also include separation of access. Systems which access registries to pull down containers should only pull from the same environment, so that unvalidated containers are not run in production and production containers are not configured in test environments. |
Design Decision 40 – Container Registries
5.11.1 Kubernetes
Kubernetes is an extension of the container model, managing the underlying hosts for containers at scale. Leveraging a cluster approach, Kubernetes offers capabilities which are essential for managing containers and applications at scale. At a high level, these capabilities include:
- Self-healing: restarting, replacing, and lifecycling containers which do not meet conditions satisfactory for the application to be “available”.
- Load-balancing: native inherent capabilities to deploy multiple instances and have them operate in parallel or sequentially.
Microsoft Azure’s iteration of Kubernetes is Azure Kubernetes Service (AKS). Microsoft offers several options for operating AKS within environments. These are described in the table below.
Configuration Options | Option 1 | Option 2 |
Network Integration | Kubenet | Azure Container Network Interface (CNI) |
Network Controls | Calico | Azure Network Policy Manager |
Management API Visibility | Private | Public |
Access Management | Native | Entra ID RBAC |
Table 13 – Azure Kubernetes Configuration Options
Microsoft updates the available AKS version frequently, in line with the overall Kubernetes kernel version release cycle.
Design Decision | |
DD-ELZ-40 | Kubernetes Kernel Version |
Decision | The minimum version of Kubernetes that will be deployed to the environment will be 1.27.1. |
Justification | Several features were progressed from ‘beta’ to ‘general availability’ with version 1.27.0. This includes several features specific for Azure, such as Azure Files support. |
Design Decision 41 – Kubernetes Kernel Version
5.11.1.1 Network integration paradigm
Kubenet is a model closer to true PaaS, where Microsoft manages the underlying networking of the cluster and components. Azure CNI is a model where networking is drawn from the deployed Virtual Network.
Design Decision | |
DD-ELZ-41 | Kubernetes Network Stack |
Decision | Kubernetes clusters will deploy the ‘Azure CNI’ network stack. |
Justification | Azure Kubernetes supports two (2) network stacks: Kubenet and Azure CNI. Azure CNI is the most appropriate, as the cluster will consume network interfaces from the underlying virtual network. This will assist in ensuring no IP configuration overlap, causing routing and communication errors. |
Design Decision 42 – Kubernetes Network Stack
5.11.1.2 Network controls tooling
Kubernetes pods, the instance of a container that runs within a deployment, have open network access by default. Azure provides two (2) methods for implementing network controls on pods: Azure Network Policy Manager or Calico.
Azure NPM is a relatively new capability which is specific to Azure’s deployment but is not yet generally available on all cluster configurations. Calico is an open-source capability which can be applied to Kubernetes agnostic to hosting environment.
Design Decision | |
DD-ELZ-42 | Kubernetes Network Controls |
Decision | Kubernetes clusters will implement Calico network policies. |
Justification | By default, Kubernetes pods may communicate at will. This is undesirable in a secure cluster, as traffic should be restricted to what is allowed and desired. |
Design Decision 43 – Kubernetes Network Controls
5.11.1.3 Management API visibility
Azure supports two (2) visibility architectures for AKS clusters: public and private. Public clusters are clusters which have been deployed so that the management API interface may be reachable over the direct Internet. Private clusters are deployed such that the management API interface can only be reached through the private network space where the cluster resides.
Design Decision | |
DD-ELZ-43 | Kubernetes Deployment Architecture |
Decision | Kubernetes clusters will deploy as ‘private’ clusters. |
Justification | Private cluster means that administrative actions to the cluster’s API server can only be performed on a system with private network access to the cluster. |
Design Decision 44 – Kubernetes Deployment Architecture
5.11.1.4 Access Management
Azure supports two (2) methods for governing access into AKS environments: native and Entra ID. Native access management means managing credentials within the AKS cluster itself for all administrative functions. Entra ID-based access management allows for the use of Entra ID identities and delegated access and privilege management with Entra ID functions.
Design Decision | |
DD-ELZ-44 | Kubernetes Access |
Decision | Kubernetes clusters will implement RBAC and restrict access to specific namespaces. |
Justification | To prevent lower privileged administrators from deploying applications to the incorrect namespace, roles-based access control can be used to isolate cluster administrators from non-cluster administrators. This ensures that admins for specific namespaces are only able to manage resources within their namespace. |
Design Decision 45 – Kubernetes Access
5.11.2 Helm
When deploying containers to Kubernetes clusters, there are a wide range of tools supporting deployments. The native Kubernetes command-line interface (CLI) tool, kubectl, can be used to apply configuration settings individually or using a configuration file.
As scale, this is not ideal, as multiple containers might be deployed with similar settings which results in repeated configuration creation and managing conflicts. This is where Helm comes into play. Helm is a packaging tool specifically for containers on Kubernetes. Alternatives to Helm include using Kubernetes manifest files (that is, not creating a “bundled” package definition) or another IaC tool such as Terraform.
Individual Kubernetes manifest files are individually created YAML files, where each specific resource type is defined in isolation from each other, even if used for the same deployment. For example, a web container would need a minimum of a service (to listen on port 80 or 443), and a deployment (to specific how many containers and what image). In a Helm chart, this would already be available and pre-defined – as a web app is the most common use-case – and just the template values passed through as parameters. Individual YAML files would have to be created, and then names and references copied to each other to work, factoring in sequence order, and changes in one would not flow through to the other.
Terraform is another IaC language that supports containers, however, is exceptionally difficult to manage containers with due to Terraform’s state management – that is, if Terraform does not push or import a change, Terraform can’t handle the change. With the underlying Kubernetes API, there are instances where the Kubernetes API will need to apply changes or alterations due to autoscaling, node management, and other cluster operations. Before a subsequent change, Terraform would need to import these changes. This would have to be a manual process, as the Kubernetes cluster appends values to the deployment which are not relevant to managing the configuration.
Considering all the above, Helm is by far the “best” option for managing container deployments at scale across clusters, as it is purpose-built for Kubernetes. It integrates and works with the Kubernetes API and only manages the configuration that the Helm chart maintainer is defining as in-scope.
Design Decision | |
DD-ELZ-45 | Container Deployment Configuration Management |
Decision | Helm will be used to manage deployment configuration(s) for containers on Kubernetes |
Justification | Helm can natively support all types of configuration deployments for containers when deploying to Kubernetes environments, with the underlying Kubernetes kernel version defining the appropriate component specification. |
Design Decision 46 – Container Deployment Configuration Management
5.11.3 Azure container tooling
5.11.3.1 Azure container registry
Azure container registry is the service used to store built Docker images, Helm charts, and Bicep modules. Images, charts, and modules can be pushed to the registry for re-use in deployments or pipelines in the environment.
5.11.3.2 Azure Kubernetes services
Azure Kubernetes Services is the service that operates Kubernetes within Azure; there are other services that provide Docker hosting. Azure operates the underlying enabling services, such as etcd, and the nodes deployed to operate Kubernetes can be used to host container-based workloads, such as Docker images or PaaS services such as Azure Functions.
5.12 Remote Access
Several mechanisms are available to provide secure remote access into an environment. In a cloud-only environment, the method for determining the most appropriate access approach is determined on whether there is a persistent need to access privately hosted systems. In most organisations, this is achieved using a network layer tool, which establishes a connection to grant access to applications and systems.
For secure environments, or where there is a need to isolate environment traffic from a broader network, this can be achieved through a virtual desktop solution, such as Remote Desktop, or Citrix. For this secure environment, the recommended approach is to leverage Azure Virtual Desktop.
Design Decision | |
DD-ELZ-46 | Remote Access Tooling |
Decision | Remote access into the environment will be provided through an Azure Virtual Desktop deployment. |
Justification | Support and administrative staff will need access into the environment at various times, particularly to view and action operational tasks. An AVD deployment will enable this capability without requiring additional tooling. |
Design Decision 47 – Remote Access Tooling
Design Decision | |
DD-ELZ-47 | Privileged Environment Access |
Decision | Privileged access within the environment will be through a separate Azure Virtual Desktop deployment. |
Justification | To support separation of duties enforcement, a separate Azure Virtual Desktop deployment is needed to isolate general activities – such as Azure DevOps work – from privileged environment support activities. |
Design Decision 48 – Restricted Environment Access
Design Decision | |
5.12.1 Azure remote access tooling
5.12.1.1 Azure virtual desktop
Azure Virtual Desktop (AVD) is an evolution of Windows Remote Desktop Services, where the configuration and infrastructure elements are hosted in Azure, and where session hosts permit a standard Windows GUI interface for a desktop. AVD can be exposed directly to the Internet or configured to only allow connections from specific private IP ranges.
6 Resiliency
Resiliency comprises multiple subject areas, not just backups. Resilient design of application workloads is essential to businesses and their customers, to ensure that they can transact and perform work.
In the context to this Solution Design, there are three (3) areas relevant to resiliency:
- Availability: the management of time where the application is available to end-users.
- Backups: the methods and copy types of persistent data volumes; and,
- Disaster Recovery: the methods and architecture by which applications can operate in critical situations.
Business Continuity is not a subject covered in this Solution Design, as the ability of the business to perform operations is determined by business processes and objectives. Disaster Recovery is the technology component represented within the broader Business Continuity Plan (BCP).
6.1 Availability
The following availability standards are well-known measurements. The given outage periods have been provided for a per-month schedule.
- 5% (2-and-a-half nines), meaning no more than 3.65 hours of unscheduled outage.
- 9% (3-nines), meaning no more than 43.83 minutes of unscheduled outage.
- 95% (3-and-a-half-nines), meaning no more than 21.92 minutes of unscheduled outage.
- 99% (4-nines), meaning no more than 4.38 minutes of unscheduled outage.
Where possible, these standards will be aligned to. Some tooling; however, has availability dictated by Microsoft due to being Platform-as-a-Service tooling, such as Entra ID.
Design Decision | |
DD-ELZ-49 | Availability Targets |
Decision | Essential Solution Design infrastructure (IaaS) services will have an availability target of 99.9% (3 nines). |
Justification | This is a stipulated requirement for the successful operation of the Azure Solution Design. |
Design Decision 50 – Availability Targets
Design Decision | |
DD-ELZ-50 | Availability Architecture |
Decision | Availability architectures will be built to support active/passive capability by default. |
Justification | Required to achieve a target of 99.9% for essential Solution Design infrastructure for production and pre-production services. This does not include Platform-as-a-Service offerings, such as Entra ID, where the availability is controlled by Microsoft. |
Design Decision 51 – Availability Architecture
6.2 Backups
Within the Azure platform, backups will be required for all data elements which are not defined as infrastructure-as-code, such as databases and virtual machines.
Azure Backup will be used to enable periodic backups of virtual machines, these backups will be managed as a backup plan and applied to all virtual machines that are individually deployed.
Design Decision | |
DD-ELZ-51 | Backup Tooling |
Decision | Azure Recovery Services Vault will be the default backup and replication tool for Azure-hosted resources. |
Justification | Azure Backup – a component of Recovery Services Vault – can support backups of all the infrastructure in use, both of virtual machines and storage volumes. |
Design Decision 52 – Backup Tooling
6.3 Disaster Recovery
Existing technology-provided Disaster Recovery principles are as follows:
- Recovery Point Objective (RPO) – 1 hour
- RPO is the maximum time tolerable for data-loss.
- Recovery Time Objective (RTO) – 24 hours
- RTO is the maximum time tolerable to return an application to operation.
The DR approach has been shaped by these availability objectives, the technical requirements of the applications, and the cross-region capabilities of the Azure persistence services.
A significant advantage to environment leveraging Infrastructure as Code (IaC) means application configurations backups and DR is greatly enhanced offering of a consistent and repeatable process to rebuild the entire environment in a timely and automated manner. It facilitates the quick rebuild of entire environments in varying regions depending on the scenario and can be validated and tested in regular intervals.
6.4 Azure resiliency tooling
6.4.1 Recovery services vault
Azure Recovery Services Vaults are a combined backup and replication tool for several Azure services. Backups are facilitated for virtual machines, storage accounts, and some compatible database services; replication is performed for virtual machines intra- or inter-region.
7 Security
Security encompasses a broad range of topics, especially when considering the wide range of capabilities and tools that can cover activities and tasks related to security. For this Solution Design, the scope of ‘Security’ is limited to the following topics or domains:
- Compliance frameworks alignment
- Security log aggregation,
- Security event and incident management
- Web application firewall
- Cloud security posture management (CSPM)
- Cloud workload protection (CWPP)
- Cloud-native application protection (CNAPP)
- Virtual machine endpoint protection (EPP)
Other items, such as authentication mechanisms or best-practice configuration settings like TLS, are not covered under this section. Instead, the Standard defined for Security Configuration should be referred to.
7.1 Compliance Framework Alignment
With any environment, the organisation for which that environment is deployed to support is subject to one or more compliance frameworks. As part of building and managing that environment, those frameworks need to be known to support implementation of secure configurations to meet those framework control points.
Design Decision | |
DD-ELZ-52 | Compliance Framework(s) Alignment |
Decision | Configure resources and services to comply with one or more compliance frameworks. |
Justification | Each environment (nonproduction, preproduction, production) and the various applications and tooling within will require compliance with different frameworks, such as ISM or CIS. Controls within resources will be configured – wherever possible – to comply with all required frameworks. |
Design Decision 53 – Compliance Framework(s) Alignment
7.1.1 Cloud security posture management
Cloud Security Posture Management (CSPM) tooling is a new capability within cloud services, to be able to report on configuration state for each individual resource and assess that configuration against a one (1) or more framework benchmarks. Framework benchmarks are created by the vendor of the CSPM tooling, or by the consumer if the tool supports custom assessments.
Whilst Azure supports this out-of-the-box with Azure Policy and Azure Defender for Cloud, the management and reporting capabilities for this are extremely immature. Several off-the-shelf items are available, however, any policy which is not available needs to be configured and build manually using Azure Policy. Azure Policy, in turn, is extremely cumbersome to manage, as each policy setting first needs to be created, then rolled-up to an Azure Policy ‘initiative’, which only then can be reported on.
Design Decision | |
DD-ELZ-53 | Cloud Security Posture Management Tooling |
Decision | Defender for Cloud will be deployed as the CSPM solution for the environment. |
Justification | The Defender for Cloud implementation is driven through toggles on subscriptions as well as implementing controls via Azure Policy configuration options. Whilst this can be effective, the configuration of Azure Policy can be extremely cumbersome to manage and maintain effectively. Azure Policy does provide some options for remediation of some settings out of the box. |
Design Decision 54 – Cloud Security Posture Management Tooling
7.2 Security Log Aggregation
Security logs form a significant basis to the ability of teams to be able to analyse, categorise, and investigate issues which pertain to security for an environment. A central log store provides a single location for those activities to occur, without needing to add additional log locations.
Azure Log Analytics, a component of Azure Monitor, is a native capability within Azure that logs can be fed into directly. It is also the service underpinning Microsoft’s Sentinel SIEM / SOAR solution.
Design Decision | |
DD-ELZ-54 | Security Log Aggregation |
Decision | Azure Log Analytics will be used to centrally capture and store security event logs, such as Entra ID audit logs or Azure activity logs. |
Justification | Azure Log Analytics is the preferred service for capturing logs from within the Azure environment, compared to raw storage blobs or event hubs. Log Analytics can structure the logs received based on their format and provide Kusto Query Language (KQL) support for searching and alerting on log events. |
Design Decision 55 – Security Log Aggregation
7.3 Security Event and Incident Management
With any environment, the ability to analyse logs at scale and provide rule-based alerting is critical for any security scenario. With Microsoft Sentinel – a modern, Azure & cloud-native event, incident, and security automation platform – leveraged over the top of Azure Log Analytics, this provides security teams and executive leadership a strong capability and tool to be able to secure and defend even the most complex cloud environments.
Design Decision | |
DD-ELZ-55 | Security Incident and Event Management Tooling |
Decision | Microsoft Sentinel will be deployed as the SIEM (and SOAR) solution for the environment. |
Justification | Microsoft Sentinel is a cloud native SIEM and SOAR solution which natively integrates into the Microsoft technology stack. As this environment is mostly deployed into Azure, Sentinel is a natural fit when augmented with additional non-Microsoft security tooling. |
Design Decision 56 – Security Incident and Event Management Tooling
7.4 Web Application Firewall
Web application firewalls (WAF) operate at the session layer of a network (aka Layer 7). These tools allow for a comprehensive understanding of the traffic that is traversing them, able to inspect, detect, and prevent malicious or undesirable activity.
There are Azure native solutions that can support this capability, along with a wide range of non-native solutions as well.
When looking at the broader security suite in operation in a given Azure environment, whilst a single underpinning detection & analysis stack is less complex architecturally and licensing-wise, this does introduce risk in that if a single component of that stack misses an event, the entire chain is likely to miss it.
The modern WAF solution requires the ability to support not just raw web / HTTP traffic, but also protect APIs, assist in discovery of configuration changes, and be managed as infrastructure-as-code to support DevSecOps and CI/CD processes.
At this point in time, a Web Application Firewall deployment is not identified as a requirement for this environment.
7.5 Cloud Workload Protection
Cloud Workload Protection Platforms (CWPP) are security tools which can secure not just traditional workloads, such as virtual machines or servers, but can also secure modern cloud workloads such as containers, PaaS storage and database platforms, and similar.
In a cloud-first or cloud-native environment, not all endpoint protection solutions can secure or supporting PaaS resources. The right CWPP solution is scalable and extensible, supporting a wide range of cloud resources, regardless of ‘as-a-Service’ archetype.
CWPP solutions can be agent-based, agentless, or a hybrid mix of both depending on components and application solution architecture.
Design Decision | |
DD-ELZ-58 | Cloud Workload Protection Tooling |
Decision | Defender for Cloud will be deployed as the CWPP solution for the environment. |
Justification | The Defender for Cloud implementation is driven through toggles on subscriptions as well as implementing controls via Azure Policy configuration options. Whilst this can be effective, the configuration of Azure Policy can be extremely cumbersome to manage and maintain effectively. Azure Policy does provide some options for remediation of some settings out of the box. |
Design Decision 59 – Cloud Workload Protection Tooling
7.6 Cloud-native Application Protection
Cloud-native Application Protection Platforms (CNAPP) are an holistic integration architecture of CSPM, CWPP, and Cloud Service Network Security (CSNS) tools -though that last one is far less common. A CNAPP solution should be able to integrate their CWPP and CSPM metrics, analyses, and reporting to provide an organisation-wide view on cloud platforms, resources, configuration, and risk profile.
Design Decision | |
DD-ELZ-59 | Cloud-Native Application Protection Tooling |
Decision | Defender for Cloud will be deployed as the CNAPP solution |
Justification | Microsoft Defender for Cloud is the new iteration of the Microsoft Defender stack for all Azure and non-Azure resources. A more-encompassing toolset than the original iteration, there is a heavy dependence on leveraging Azure Policy for control. |
Design Decision 60 – Cloud-native Application Protection Tooling
7.7 Virtual Machine Endpoint Protection
The most familiar security tool to end-users and IT professionals alike, endpoint protection (EPP) is the all-encompassing term for ‘next-generation’ anti-virus (NGAV) and enhanced-detection-and-response (EDR) tooling. Compared to traditional ‘signature’ based anti-virus solutions, EPP solutions deploy agents to endpoints, to collect and stream activity logs from supported endpoints – usually virtual machines or servers – to build a pattern of operational behaviour to determine threats. This is enhanced through the use of shared threat intelligence which can be updated on the fly in the back-end of the EPP vendor, allowing for rapid iteration and detection of new and emerging threats without the endpoint needing to receive updates or ‘definitions’.
Design Decision | |
DD-ELZ-60 | Virtual Machine Endpoint Protection Tooling |
Decision | Microsoft Defender for Endpoint will be deployed as the EDR and NGAV solution for virtual machines. |
Justification | Defender for Endpoint is a native service to Azure, supporting the intended Operating Systems of virtual machines within the environment. |
Design Decision 61 – Virtual Machine Endpoint Protection Tooling
7.8 Vulnerability Management
Vulnerability management is a process of identifying, evaluating, prioritising, and mitigating security vulnerabilities in an organisation’s IT infrastructure, systems, and applications. The goal of implementing a vulnerability management program is to assist in proactively prevent security breaches and protect critical assets by addressing vulnerabilities before they can be exploited.
To achieve this, a vulnerability assessment capability is used to support the vulnerability management process. Vulnerability assessment solutions identify and evaluate workload and/or network vulnerabilities by constantly scanning and monitoring the solution attack surface for risks. With the use of cloud-native, public cloud platforms that are dynamic, components of the assessment workflow differ significantly from vulnerability assessments for on-premises datacentres, due to the split-responsibility model with public cloud.
Solutions that support both may overlap, however, solutions built for traditional infrastructure are not as mature for cloud-native offerings. In the inverse, solutions built for cloud-native workloads typically ignore or are not supported for traditional Infrastructure-as-a-Service capabilities.
Design Decision | |
DD-ELZ-61 | Vulnerability Management |
Decision | Vulnerability assessment of containers will leverage Defender for Endpoint and Cloud |
Justification | The Microsoft Defender suite is an over-arching solution which supports container-native approaches and integrates vulnerability scanning into Kubernetes runtime protection effectively. |
Design Decision 62 – Vulnerability management (containers)
Design Decision | |
DD-ELZ-62 | Vulnerability Management |
Decision | Vulnerability assessment of IaaS VMs will leverage Defender for Cloud |
Justification | Defender for Cloud is a proven vulnerability assessment solution on traditional IaaS components, such as virtual machines and network infrastructure. Use of Microsoft Defender across both Kubernetes and IaaS will present a single view of vulnerability detection and management. |
Design Decision 63 – Vulnerability management (IaaS VMs)
8 Operations
8.1 Logs and Metrics
Metrics for Azure hosted resources will be collected in Azure Monitor. These will be used to trigger alerts. All logs for services running inside the Azure platform will be shipped to Azure Monitor. This will provide for tamper-proof, near real-time auditing.
Design Decision | |
DD-ELZ-63 | Monitoring Tooling |
Decision | Azure Monitor will be used as the default monitoring solution for the environment. |
Justification | The Azure Monitor stack, including Log Analytics and Alerts, underpins a significant volume of the logging and metrics capability within Azure. Azure Monitor supports each Azure service natively, to varying degrees. |
Design Decision 64 – Monitoring Tooling
Design Decision | |
DD-ELZ-64 | Operational Log Aggregation |
Decision | Azure Log Analytics will be used to centrally capture and store operational logs and metrics, such as Container Insights, for platform management services. |
Justification | Azure Log Analytics is the preferred service for capturing logs from within the Azure environment, compared to raw storage blobs or event hubs. Log Analytics can structure the logs received based on their format and provide Kusto Query Language (KQL) support for searching and alerting on log events. |
Design Decision 65 – Operational Log Aggregation
8.2 Alerting
Azure Monitor will be utilised for alerts relating to the Azure infrastructure. Alerts will be configured to include a webhook call (or similar) to an ITSM solution, such as ServiceNow. This will minimise the time taken for these alerts to be issued and tracked in ITSM ticketing solutions.
This eliminates Azure Monitor as a single point of failure in the alerting chain. Additionally, it provides more contextual information about each alert and direct links back to the source monitoring system to investigate and resolve the issue.
8.3 Update Management
Update Management refers to the practice and process by which security and software updates are applied to workloads. Within the Azure environment, the following workload types require a configuration for managing updates:
- Azure Virtual Desktop Session Hosts
- Virtual Machines
- Container Images
- Azure Kubernetes Services
There are several tools suitable for use within Azure which support different capabilities for managing updates, each of which have their own limitations and caveats. Several of the current generally available services are due to be decommissioned in 2024, as Azure phases out support of the legacy Azure Log Analytics agent.
Design Decision | |
DD-ELZ-65 | Update management for Azure Virtual Desktop Session Hosts |
Decision | Azure Virtual Desktop Session Hosts will be updated through a release management process |
Justification | Azure Virtual Desktop VMs runs a client operating system and software, not server software, to be able to manage the environment. As the release of Windows patches is on a known calendar, new images can be created automatically and updated into AVD instead of patching VMs in place. |
Design Decision 66 – Update management for Azure Virtual Desktop Session Hosts
Design Decision | |
DD-ELZ-66 | Update management for Virtual Machines |
Decision | Azure Update Management will be used to manage updates on Virtual Machines |
Justification | Azure Update Management (Preview), whilst still being in preview, supports the integration of generalised VM images pushed to Azure Compute Gallery. As each update will need to be scheduled in a patching calendar, the use of ‘Automatic VM Updates’ is a function that is not presently required. |
Design Decision 67 – Update management for Virtual Machines
Design Decision | |
DD-ELZ-67 | Update management for Container images |
Decision | Container images will need to be updated in their build configuration |
Justification | Containers are immutable infrastructure, in that their base configuration cannot be updated after build. Instead, Dockerfiles will need to be updated and rebuilt, and then the new image deployed to Kubernetes. |
Design Decision 68 – Update management for Container images
Design Decision | |
DD-ELZ-68 | Update management for Azure Kubernetes Services |
Decision | Azure Kubernetes Services will set to update automatically within the cluster configuration. |
Justification | AKS supports several options for managing versions on nodes. Azure pushes updates to Linux nodes daily, and Windows nodes on a manual schedule. Through the cluster configuration, the cluster and node kernel versions and statuses will be updated in parallel. |
Design Decision 69 – Update management for Azure Kubernetes Services
Document | Reference |
Microsoft Cloud Adoption Framework | https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ |
Microsoft Well-Architected Framework | https://learn.microsoft.com/en-us/azure/well-architected/ |
Kubernetes | https://kubernetes.io/docs/home/ |
Australian Cyber Security Centre Information Security Manual | https://www.cyber.gov.au/resources-business-and-government/essential-cyber-security/ism |
Centre for Information Security | https://www.cisecurity.org/ |
National Institute of Standards & Technology Cybersecurity Framework | https://www.nist.gov/cyberframework |
National Institute of Standards & Technology Security and Privacy Controls for Information Systems and Organizations | https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final |
International Standards Organisation | https://www.iso.org/obp/ui/#iso:std:iso-iec:27001:ed-3:v1:en |
Tenant Hierarchy
Azure Cloud Platform |
Contents
2.1 Conceptual Architecture. 5
3 Azure Governance Hierarchy. 6
5 Monitoring Customer Experience. 11
Appendix A Enterprise Scale Landing Zone Components. 12
Appendix B Conceptual ESLZ Structure. 13
Tables
Table 2 – RLX Contact Details. 2
Table 3 – RLX Version Control Details. 2
Table 4 Design Requirements. 4
Table 6 Subscription Groups. 8
Figures
Figure 1 Conceptual Organisation. 5
Figure 3 Subscription Structure. 8
Figure 4 Subscription Groups. 9
Figure 5 Conceptual ESLZ Architecture. 13
9 Overview
9.1 Purpose
This document provides the detailed design for the Management and Subscription Group deployment as stipulated in the Enterprise Scale Landing Zone Solution Design document.
9.2 In-scope
The following components of the platform are in scope for this document:
- Management Group and Subscription Organisation.
- Naming and Tagging convention.
9.3 Design Requirements
The following are the design requirements inherited from the high-level solution design.
ID | Design Decision | Rationale |
DD-ELZ-2 | Only a single tenant will be deployed. | Whilst multiple tenants could separate Production from Nonproduction at the authentication step, given the relatively small size of the organisation, this may represent too much overhead and complexity to manage. Subscriptions can be used instead to separate those environments, along with the use of Privilege Management for temporary elevated Production access. |
DD-ELZ-3 | Management groups will be used to group subscriptions together logically. | Management groups allow for the hierarchy of subscriptions and management groups. This means that access policies, Azure policies, and the like to be applied at different levels. For example, global ‘Reader’ access applied at the parent management group allows for all child objects to inherit that assignment. |
DD-ELZ-4 | Subscriptions will be used to segregate resources per functional domain and environment. | Using multiple subscriptions, aligned to business structure, allows for a separation of cost breakdown far more easily. Owners can be assigned view access to their subscriptions to review cost, and not risk confusing costs that are owned by other owners. |
DD-ELZ-5 | The default language of the tenant will be set to ‘English’. | The geographical region this environment is to be deployed to is Australia. The national language of Australia is English. |
DD-ELZ-7 | Australia East will be the Azure region used for hosting production environments. | Australia East represents the Azure region with the greatest capacity and availability of services; all of which are rated to the PROTECTED level. |
DD-ELZ-8 | Australia Southeast will be the Azure region used for hosting DR environments. | Australia Southeast is the paired region for Australia East. |
DD-ELZ-9 | Australia East will be the Azure region used for hosting nonproduction environments. | Australia East represents the Azure region with the greatest capacity and availability of services; all of which are rated to the PROTECTED level. |
10 Solution Architecture
The CUST Tenancy is the foundational infrastructure capability provider, from which CUST can design, build, deploy, and manage the full spectrum of applications required to support all business technology initiatives.
10.1 Conceptual Architecture
The conceptual state hosting platform is the Azure Cloud. Azure IaaS and PaaS services will be used to
provide the platform level support services.
Figure 1 Conceptual Organisation
11 Azure Governance Hierarchy
11.1 Management Groups
Management Groups are used to provide enterprise grade management at scale, providing a flexible hierarchy for unified policy and access management.
The following table outlines the Management Groups that will be deployed for CUST.
Management Group ID | Management Group Display Name | Parent | Function |
Tenant Root Group | Tenant Root Group | – | The default root management group will be left with existing structures remaining unchanged within the current Microsoft Azure tenant. This will also allow future changes to be incorporated to accommodate expansion of management groups as needed. |
mg-root | Root | Tenant Root Group | The Azure tenant’s top-level management group serves as the central container for customer role definitions, custom policy definitions, and Microsoft global policy assignments. However, it has limited direct role assignments. At this scope, the primary objective of policy assignments is to uphold security and autonomy for the platform, particularly as additional child management groups and subscriptions are established. |
mg-platform | Platform | Mg-root | The Platform management group will serve as the parent for two child management groups: Production and Non-Production. These child management groups will further include Connectivity, Management, Integration, and Identity child management groups. At the Platform management group level, RBAC permissions and Azure policies will be assigned to roles responsible for overseeing and maintaining the platform. Additionally, key Azure Policies will be implemented to guide the construction of the platform. |
mg-management | Management | Platform | The Management subscription, residing within this management group, will cater to centrally managed platform infrastructure. It will provide comprehensive and scalable management capabilities across all landing zones within the Microsoft Azure platform. |
mg-connectivity | Connectivity | Platform | To establish end-to-end connectivity for all landing zones within the Microsoft Azure platform, a dedicated subscription will be allocated for centrally managed networking infrastructure. This subscription will include the deployment of virtual network hubs, ExpressRoute circuits, associated gateways, and firewalls. |
mg-identity | Identity | Platform | This management group contains a dedicated subscription for identity. This subscription is a placeholder for Microsoft Entra Domain Services. The subscription also enables AuthN or AuthZ for workloads within the landing zones. Specific Azure policies are assigned to harden and manage the resources in the identity subscription. |
mg-security | Security | Root | For hosting security related services, a dedicated security subscription will be designated. Ensures segregation and centralization of tenant wide security tooling. |
mg-applications | Applications | Root | Workloads will be provisioned in subscriptions that reside within child management groups of the Applications management groups. This hierarchical structure enables a flexible and targeted approach to policy assignments, ensuring clear separation between active Applications and Sandpit or Decommissioned subscriptions. |
mg-jira | Jira | Applications | Used to host and provision Jira specific subscriptions and workloads. |
mg-sandpit | Sandpit | Root | For application teams and individual users seeking to experiment and evaluate Azure services, dedicated subscriptions will be established within the Sandpit management group. This group is designed with specific policies to prevent any control plane or data plane access to production environments. This includes restricting network connectivity to the Connectivity Landing Zones and, subsequently, the on-premises environment. |
11.2 Subscriptions
Subscriptions provide the primary administrative, security and billing boundaries. Every Azure resource (VM, network, managed database, storage account etc.) belongs to a Subscription. The Management and connectivity subscriptions will be segregated into non-production and production.
Figure 3 Subscription Structure
Subscription Display Name | Parent |
CUSTPD-Security | mg-Security |
CUSTPD-Identity | mg-Identity |
CUSTPD-Integration | mg-Integration |
CUSTNP-Integration | mg-Integration |
CUSTPD-Connectivity | mg-Connectivity |
CUSTNP- Connectivity | mg-Connectivity |
CUSTPD-Management | mg-Management |
CUSTNP- Management | mg-Management |
CUSTPD-Jira | mg-Jira |
CUSTNP-Jira | mg-Jira |
CUSTNP-Sandpit | mg-Sandpit |
12 Resource Locks
Resource locks are an additional Azure control that can be applied to prevent the alteration or deletion of critical resources and infrastructure in the environment. Adding resource locks adds an additional degree of defence to prevent users with malicious intent or from making unintentional mistakes.
There are currently two supported lock types:
- Read-Only: Allows read-only access to resources but prevents modification and deletion.
- Delete: Allowed authorised modification but prevents deletion.
Resource locks will be applied at a range of different levels across the enterprise scale landing zone. This includes at subscription, resource group, resource levels, which will be defined in the As-Built documentation at build completion.
13 Monitoring Customer Experience
Customer Experience is one of RLX’s core values. Meeting and exceeding our customers’ expectations is a goal we take very seriously. Our aim is to be a trusted partner, solving our customer’s cyber security challenges with intelligence led solutions that are outcome focused and operationally informed.
To ensure we are meeting this goal and continuously improving our customers’ experience, RLX has implemented a series of formal and informal measures to regularly measure customer satisfaction. These include identifying key points of contact, regularly scheduled meetings, and a formal Voice of the Customer platform utilising the industry best practice Net Promoter Score (NPS). NPS measurement is applied across all areas of RLX and focuses on two key components of customer satisfaction: transactional (related to a particular project or interaction) and relationship (the overall relationship between RLX and our customers across multiple stakeholders).
As a RLX customer, you may from time to time receive a short survey requesting your feedback on us as an organisation and the services we provide. Surveys will take no longer than a few minutes to complete. All results are regularly reviewed by RLX’s leadership, along with any subsequent improvement activities undertaken as a result. While there is an option to unsubscribe if you prefer, we appreciate your honest feedback and thank you in advance for taking the time to respond.
DESIGN AREA | DESCRIPTION |
Enterprise agreement and AAD tenant | An Enterprise Agreement enrolment represents the commercial relationship between Microsoft and CUST. An Entra ID tenant provides identity and access management, which is an important part of your security posture. In an Entra ID tenant, authenticated and authorized users have access only to the resources for which they have access permissions. |
Identity and access management | Build Entra ID design and integration to ensure both server and user authentication and role-based access control (Azure RBAC). |
Management group and subscription organisation | Subscriptions – the primary administrative, security and billing boundaries. Every Azure resource (VM, network, managed database, storage account etc.) belongs to a Subscription. Management Groups that define a grouping hierarchy on which to apply security and policies that are inherited to the subscriptions within the management group. |
Management and monitoring | Design, deploy, and integrate platform-level holistic (horizontal) resource monitoring and alerting. Define and streamline operational tasks, such as patching and backup. |
Network topology and connectivity | To ensure north-south and east-west connectivity between platform deployments, build and deploy the end-to-end network topology across Azure regions and on-premises environments |
Business continuity and disaster recovery | Make sure corporate, regulatory, and line-of-business controls (not defined in entirety yet) are in place. Identify, describe, build, and deploy holistic and landing-zone-specific policies to place those controls |
Security governance and compliance | Use policies to guarantee the compliance of applications and underlying resources without any abstraction provisioning or administration capability. |
Platform automation and DevOps | Ensure a safe, repeatable, and consistent delivery of infrastructure-as-code artifacts. Design, build, and deploy an end-to-end DevOps experience with robust software development lifecycle practices to ensure artifact delivery. |
Figure 5 Conceptual ESLZ Architecture
Naming Standard
Standard
Azure Cloud Platform |
Contents
3.2 Privileged user accounts. 7
3.2.1 Administrative privileged accounts. 7
3.2.2 Security privileged accounts. 7
3.3 Emergency access accounts. 8
3.4 Enterprise Applications. 8
3.7 Conditional Access Locations. 9
3.8 Conditional Access Policies. 10
5 Monitoring Customer Experience. 16
Tables
Table 2 – RLX Contact Details. 2
Table 3 – RLX Version Control Details. 2
Table 4 – RLX Version Control Changelog. 2
Figures
No table of figures entries found.
14 Overview
This document is intended for use to guide the creation and configuration of resources within Azure and Entra ID for this environment. Specifically, the use of a naming standard or convention to be able to identify the scope at which a resource is used for.
Note that throughout this document, upper-case letters have been used to identify characters used within names. When actually deployed to Azure or Entra ID, all names should be deployed using lower case.
15 Environment Names
This Azure environment will host several different environments, which will be used to develop, support, and operate application workloads. The below table describes the convention by which each environment will be named.
15.1 Tenant ID
Each tenant is assigned a 3-letter identifier. The tenant identifier for this environment is:
- Australian CUST Corporation (CUST)
15.2 Environments
There are three (3) discrete environments, namely:
Name | Abbreviation |
Production | PD |
Pre-production | RE |
Nonproduction | NP |
These environments are isolated from each other and may be comprised of one or more tiers.
15.3 Tiers
The following table describes the individual tiers, or ‘sub-environments’, within each named Environment.
Tier Name | Abbreviation | Parent Environment |
Production | PRD | Production |
Disaster Recovery | DRP | Production |
Staging Production | STP | Pre-production |
Staging Disaster Recovery | STD | Pre-production |
Training | TRN | Nonproduction |
User Evaluation & Testing | UET | Nonproduction |
Integration | ITF | Nonproduction |
Development / Development 2 | DEV | Nonproduction |
Sandbox / Development 1 | SBX | Nonproduction |
16 Azure Active Directory
16.1 User accounts
User accounts are unprivileged accounts which are used for day-to-day activities, such as accessing email or communicating with partner organisations.
User accounts are to be created with the following standard:
givenname.surname@support.CUST.au
Where multiple users share identical given names and surnames, eg two (2) John Smith’s, the following standard will apply:
First user: givenname.surname@support.CUST.au
Second user: givenname.surname1@support.CUST.au
The integer identifier will increment by 1 for each subsequent user of the same given name and surname.
16.2 Privileged user accounts
Privileged accounts are accounts used when performing work in a technical system which has permissions or roles assigned which allow for reading, changing, or otherwise editing configuration settings.
16.2.1 Administrative privileged accounts
Privileged user accounts are to be created with the following standard:
adm-<user account name>@CUST.onmicrosoft.com
The use of the Microsoft tenant name, rather than the custom domain name, is to assist in identification and to reduce risk in the event of issues with the custom domain name.
16.2.2 Security privileged accounts
Privileged user accounts are to be created with the following standard:
sec-<user account name>@support.CUST.au
16.3 Emergency access accounts
Emergency access, or so-called “break-glass” accounts are privileged accounts which are used in the event of an emergency, when ‘normal’ access through user and/or privileged accounts is not feasible.
Rather than define a convention, the following table describes the five (5) emergency access accounts to be created.
Additionally, each of these accounts has a physical multi-factor authentication token attached to each of them.
Name | Azure / Entra ID Role(s) | Function |
TBC | Global Administrator | Emergency access to Azure Active Directory. |
TBC | Global Administrator | Emergency access to Azure Active Directory. |
16.4 Enterprise Applications
Enterprise Applications are Entra ID objects used for managing user access to applications. This could be an integrated single-sign-on application, such as Microsoft Office, or custom applications built for a tenant.
The following table describes the guidance attached to Enterprise Applications:
Function | Guidance | Example |
Web Application | Use the FQDN for the name | ServiceNow application: |
Management Tool | Use the product name for the name | F5 Distributed Cloud: |
16.5 Service Principals
Service Principals are Entra ID objects used for configuring application authentication to Entra ID. This could be used for an application to support SSO with an Enterprise Application or a client application used to execute privileged tasks within Azure.
They can also be known as “App Registrations”. Service principals are to be created with the following standard:
<Identifier>-<Application Short Name>-<Party>-<Function>
Name Component | Description | Example |
Identifier | Identifier of object type | SPN (Service Principal) |
Application Short Name | A short-form version of the application’s name | GHE (GitHub Enterprise) |
Party | The partner or vendor who uses the Service Principal | CS2 (RLX CS2) |
Function | What function is performed using the Service Principal | ACTION (GitHub Actions) |
16.6 Groups
Groups in Entra ID are used to logically group items for access, roles, or similar. Groups can be used to manage access to roles, applications, or resources within Azure.
Groups are to be created with the following standard:
<Identifier>-<Area>-<Subject>-<Role>
Name Component | Description | Example |
Identifier | Identifier of object type | Group: – GP |
Area | The area where the group will be used | Azure: – AZR Entra ID: – EID |
Usage | Used for users or privileged access | User membership: – USR Privileged access: – PAG |
Subject | The application, role, or subscription the group is to be used for. | Azure: – Subscription name – Management group Entra ID: – Application – Entra ID role |
Role | The role or level of privilege associated with the group. | Azure: – Contributor – Reader Entra ID: – Global Reader – Global Administrator |
16.7 Conditional Access Locations
To avoid confusion, the use of a strict location naming policy is recommended. While something like Australia is obvious, “Melbourne Office” can become confusing where there are multiple matching locations or multiple network providers. The below are recommendations but should be fine-tuned, relevant to the organisation.
Conditional Access locations are to be created with the following standards:
Corporate locations: <Country Code>-<City Code>-<Type>-<Address Summary>
Provider locations: SVCLOC-<Provider>-<Usage>
Name Component | Description | Example |
Country Code | 2-letter code | Country code: – AU (Australia) – FR (France) – GB (United Kingdom) |
Identifier | Identifier of location type | SVCLOC |
City Code | 3-letter code for the city | City code: – MEL (Melbourne, AU) – ADL (Adelaide, AU) – SYD (Sydney, AU) |
Type | Facility type | Office, Datacentre |
Address | Notable name | Address: – SOC Office |
Provider | Service Provider name | Provider: – AWS – Microsoft – Okta |
Usage | Function of location | Location function: – Marketing – Admin – SOC |
16.8 Conditional Access Policies
For managing two or three basic policies, organisations can avoid a strict naming policy. However, when you start managing different use case scenarios, whether that be groups of users, applications, or controls, add in multiple administrators and continued growth, it will rapidly become an unmanageable beast. Any naming policy should use the following key attributes:
- A reference number for change management; and,
- The application(s) in scope; and,
- The controls enforced; and,
- Who it applies to; and,
- When it applies
Conditional Access policies are to be created with the following standard:
<Policy Number> – <User Scope> – <Applications> – <Conditions> – <Action> – <Control>
Name Component | Description | Example |
Policy Number | A 3-digit integer counter | Number: – 100 – 120 – 300 |
User Scope | Type of scope for policy | Scope: – All Users – Administrators – Guests |
Applications | Applications under scope of the policy | Application: – All Applications – Microsoft 365 – Selected |
Conditions | Conditions attached to the policy | Condition types: – Location – Device Type – Sign-in Risk |
Action | The action to take under the policy | Action: – Allow / Grant – Deny / Block |
Control | The control enforcing the policy | Control: – Sign-in frequency – Multi-factor Authentication |
17 Azure Resources
When being deployed to Azure, each resource type needs to be identifiable, to understand:
- What type of resource it is
- Where the resource is deployed
- What environment or environment tier it is deployed to
This is essential to understand the number of resources are in use and how they are consumed. Names, when combined with tags, are used to produce a comprehensive breakdown of what “assets” are present within the environment.
Azure resources are to be created with the following standard, based on permitted characters:
Name Type | Name Structure |
Standard | <Identifier>-<Subscription>-<Region>-<Environment Tier>-<Function>-<Integer> (The Integer is now optional and preferably removed) |
Management groups | <Identifier>-<Function> |
Key Vaults | kv-<Subscription>-<Region>-<Environment Tier>-<Function>-<Integer> – Integer – Region Reason: Key Vaults are limited to 24 characters |
Storage Accounts | sa<Subscription><Region><Environment Tier><Function><integer> Integers are limited to 1 or 2 digits, not 3. |
Azure Policy Initiatives | in-<Initiative Name> |
Azure Policy Policies | py-<Initiative>-<Policy Name> |
Private Endpoints | pe-<Resource Name>-<Integer> |
Network Interfaces | nic-<VM Name>-<Integer |
Network Security Groups | nsg-<VNet Name>-<Subnet Name> |
Subnets | sn-<Function> |
VNet Peering | Peer-<Hub Vnet Name>-<Spoke Vnet Name> |
Parameters for Azure resources are described below.
Name Component | Description | Example |
Identifier | The resource type unique identifier | Identifier: – kv (Key Vault) – sa (Storage Account) |
Subscription | The short-form name of the subscription | Subscription: – con (Connectivity) – mgm (Management) |
Region | The short form name of the Azure region | Region: – ae (Australia East) – as (Australia Southeast) |
Environment Tier | The 3-letter ID of the environment tier | Environment Tier: – prd (Production) – trn (Training) |
Function | A 5-letter function for the resources | Function: – eslz (Enterprise Landing Zone) |
Integer | A 3-digit integer for the resource | Integer: – 001 |
17.1 Reserved Words
The following table contains the words which are reserved for use and cannot be used to create Azure resources.
Reserved Names | |||
Access | Azure | Bing | Bizspark |
Biztalk | Cortana | DirectX | Dotnet |
Dynamics | Excel | Exchange | Forefront |
Groove | Hololens | Hyperv | Kinect |
Lync | MSDN | O365 | Office |
Office365 | OneDrive | OneNote | Outlook |
PowerPoint | SharePoint | Skype | Visio |
VisualStudio | Login | Microsoft | Windows |
Xbox | |||
17.2 Regions
Azure operates in multiple geographical regions as well as some resources which operate region-agnostically. The following table describes the list of regions in use with this environment and the ‘short form’ name each region will be assigned in names.
Azure Regions | Abbreviation |
Global | GL |
Australia Southeast | AS |
Australia East | AE |
17.3 Management
Resource | Abbreviation | Examples |
Management group(s) | MG | mg-tenant |
Subscription(s) | #N/A | pbpd-connectivity |
Resource group(s) | RG | rg-pb-con-ae-prd-eslz-001 |
17.4 Policy
Resource | Abbreviation | Examples |
Initiative(s) | IN | in-iso27001 |
Policy(ies) | PY | py-iso27001-kv |
17.5 Networking
Resource | Abbreviation | Examples |
Application Gateway(s) | AGW | agw-con-ae-prd-eslz-001 |
Application Security Group(s) | ASG | asg-con-ae-prd-avd-001 |
ExpressRoute Circuit(s) | ERC | erc-con-ae-prd-eslz-001 |
ExpressRoute Connection(s) | CON | con-con-ae-prd-eslz-001 |
Firewall(s) | FW | fw-con-ae-prd-eslz-001 |
Firewall Policy(ies) | FWP | fwp-con-ae-prd-eslz-001 |
IPSec Connection(s) | IPSEC | ipsec-con-ae-prd-eslz-001 |
Load Balancer(s) – Internal | ILB | ilb-con-ae-prd-eslz-001 |
Load Balancer(s) – Public | ELB | elb-con-ae-prd-eslz-001 |
Local Network Gateway(s) | LNG | lng-con-ae-prd-eslz-001 |
NAT Gateway(s) | NGW | ngw-con-ae-prd-eslz-001 |
Network Interface(s) | NIC | nic-<vm name>-001 |
Network Security Group(s) | NSG | nsg-<vnet name>-<subnet name> |
Private Endpoint(s) | PE | pe-<resource name>-001 |
Public IP Address(es) | PIP | pip-con-ae-prd-eslz-001 |
Public IP Prefix(es) | PFX | pfx-con-ae-prd-eslz-001 |
Route Table(s) | RT | rt-<vnet-name>-<function> |
Virtual Network(s) | VN | vn-con-ae-prd-eslz-001 |
Virtual Network Gateway(s) | VNG | vng-con-ae-prd-eslz-001 |
Virtual Network Peering(s) | PEER | peer-<vnet name>-<spoke/hub> |
17.6 Compute
Resource | Abbreviation | Examples |
Availability Set(s) | AS | as-mgm-ae-prd-avd-001 |
Disk(s) – OS | OSD | disk-<vm name>-osd |
Disk(s) – Data | DATADISK | disk-<vm name>-datadisk-001 |
Image Gallery(ies) | ACG | acgpbpd001 |
Virtual Machine(s) – Linux | LX | lxmgmasdevodb01 |
Virtual Machine(s) – Windows | WD | wdmgmasdevado01 |
Virtual Machine(s) – AVD | #N/A | deasep11608 |
17.7 Azure Virtual Desktop
Resource | Abbreviation | Examples |
Host Pool(s) | HP | hp-mgm-ae-prd-avd-001 |
Scaling Plan(s) | SP | sp-mgm-ae-prd-avd-001 |
Application Group(s) | AG | ag-mgm-ae-prd-avd-001 |
Workspace(s) | WS | ws-mgm-ae-prd-avd-001 |
17.8 Modern Apps
Resource | Abbreviation | Examples |
Application Configuration Store(s) | ACS | acs-sec-ae-prd-vms-001 |
App Service Plan(s) | ASP | asp-sec-ae-prd-vms-001 |
Function App(s) | FA | fa-sec-ae-prd-vms-001 |
Notification Hub Namespace(s) | NH | nh-sec-ae-prd-vms-001 |
17.9 Secrets Management
Resource | Abbreviation | Examples |
Key Vault(s) | KV | kv-mgm-ae-prd-iac-001 |
Hardware Security Module(s) | HSM | hsm-mgm-ae-prd-iac-001 |
17.10 Storage
Resource | Abbreviation | Examples |
Storage Account(s) | SA | samgmaeprdiac001 |
17.11 NetApp Files
Resource | Abbreviation | Examples |
Account(s) | NAF | naf-sbx-ae-npd-db-001 |
Capacity Pool(s) | NCP | ncp-sbx-ae-npd-db-001 |
Volume(s) | NVL | nvl-sbx-ae-npd-db-001 |
Security Standard
Azure Cloud Platform |
Contents
w Asymmetric cryptographic algorithms. 5
w Asymmetric cryptographic key exchanges. 6
w Symmetric cryptographic algorithms. 6
w Symmetric cryptographic ciphers. 7
o Hash-based message authentication codes. 7
w Internet Protocol Security. 9
w Certification Authorities. 11
o Root Certification Authorities. 11
o Intermediate Certification Authorities. 11
o Virtual machine image updates. 12
o Virtual machine update management. 12
w Monitoring Customer Experience. 14
Appendix A Document References. 15
Tables
Table 2 – RLX Contact Details. 2
Table 3 – RLX Version Control Details. 2
Table 4 – RLX Version Control Changelog. 2
Table 5 – Asymmetric encryption comparison in bit-size. 5
Table 6 – Preference order for asymmetric encryption algorithms. 5
Table 7 – Preference order for asymmetric key exchange algorithms. 6
Table 8 – Preference order for symmetric encryption algorithms. 6
Table 9 – Preference order for symmetric ciphers. 7
Table 10 – Hash algorithm preference order. 7
Table 11 – HMAC preference order. 8
Table 12 – TLS 1.3 ciphersuite preference order. 8
Table 13 – TLS 1.2 ciphersuite preference order. 9
Table 14 – IPSec IKE Phase 1 minimum settings. 10
Table 15 – IPSec IKE Phase 2 minimum settings. 10
Table 18 – Virtual machine image update schedule. 12
Table 19 – Virtual machine update schedule. 12
Figures
No table of figures entries found.
This Standard document is intended to act as the guiding document when deploying Cloud resources and configurations. Whilst that is not an exhaustive scope item, this should cover most use cases within the Azure platform.
Asymmetric cryptography, also known as public key cryptography or public/private keypairs, is a form of cryptography where different keys are used to encrypt and decrypt data. The public key can only be used to encrypt data and the private key is used to decrypt data.
This type of cryptography is most frequently encountered when browsing to a website over HTTPS, also known as SSL or TLS, or when using public key authentication for Secure Shell (SSH) and Secure Shell File Transfer (SFTP).
There are many implementations of asymmetric cryptography in use, determined based on the method of mathematical computation of the private key. Each of these algorithms can generate keys of different sizes, where smaller keys are both less intensive to generate and easier to permutate. Algorithms are most frequently used to generate X.509 certificates, which serve as digital signatures, for a variety of functions such as HTTPS certificates.
The following table shows a key size comparison between RSA and Elliptic-Curve (ECDSA) algorithms.
Security (bit-size) | RSA key size to meet Security | ECDSA key size to meet Security |
80 | 1024 | 160 |
112 | 2048 | 224 |
128 | 3072 | 256 |
192 | 7680 | 384 |
256 | 15360 | 512 |
Table 5 – Asymmetric encryption comparison in bit-size
The following combination of algorithms and bit sizes are listed in descending preference order, as a recommendation in-line with the Australia Cyber Security Centre’s guidance on asymmetric cryptography.
Preference Order | Name | Algorithm | Key Bit Size |
1 | ECDSA-521 | secp521r1 | 521 |
2 | ECDSA-384 | secp384r1 | 384 |
3 | ECDSA-256 | secp256r1 | 256 |
4 | RSA-4096 | RSA | 4096 |
5 | RSA-3072 | RSA | 3072 |
Table 6 – Preference order for asymmetric encryption algorithms
When two systems communicate, or desire to, through an encrypted channel, these systems must perform a ‘handshake’ where the two systems agree on a shared set of elements – an algorithm, a cipher, and a message authentication code. As part of this handshake, the two systems perform an exchange of public keys to agree on a single key for the duration of the channel.
The following key exchanges are approved for use, listed in descending preference order, as a recommendation in-line with the Australia Cyber Security Centre’s guidance on asymmetric cryptography.
Preference Order | Acronym | Name |
1 | ECDHE | Elliptic-Curve Ephemeral Diffie-Helman |
2 | DHE | Ephemeral Diffie-Helman |
Table 7 – Preference order for asymmetric key exchange algorithms
Symmetric cryptography is a form of cryptography where the same key is used to both encrypt and decrypt data.
There are many implementations of symmetric cryptography in use, determined based on the method of mathematical computation of the private key. Each of these algorithms can generate keys of different sizes, where smaller keys are both less intensive to generate and easier to permutate. The most common symmetric algorithm in use today is AES (Advanced Encryption Standard).
The following combination of algorithms and bit sizes are listed in descending preference order, as a recommendation in-line with the Australia Cyber Security Centre’s guidance on symmetric cryptography.
Preference Order | Name | Algorithm | Key Bit Size |
1 | AES-256 | AES | 256 |
2 | AES-192 | AES | 192 |
3 | AES-128 | AES | 128 |
Table 8 – Preference order for symmetric encryption algorithms
In this context, ciphers are the method by which data is encrypted either as a block or a stream of data. These are respectively known as block ciphers and stream ciphers.
The following symmetric ciphers are listed in descending preference order, as a recommendation in-line with the Australian Cyber Security Centre’s guidance on symmetric cryptography.
Preference Order | Acronym | Name | Cipher Type |
1 | GCM | Galois/Counter-Mode | Block-to-Stream |
2 | CBC | Chain Block Cipher | Block Cipher |
Table 9 – Preference order for symmetric ciphers
A hash function, hash or hashing, is a method or algorithm which transforms a block of data of a given size to a hexadecimal string. Hash functions operate one-way; that is, performing a hash of content cannot be ‘unhashed’. Rather, using a hash algorithm on the same content always results in the same hash output. As an example, the string “fox” might be translated to “ae5512b04cf” through a hash function.
Hash functions are frequently used when storing passwords, along with a pseudo-random component known as a ‘salt’ which enables the resulting ‘hashed’ content to be different, even with the same source content. This also increases the computational power required to perform the hashing process, which in turn increases the security of the end-result.
The following list of hashing functions are listed in descending preference order, as a recommendation in-line with the Australia Cyber Security Centre’s guidance on asymmetric cryptography.
Preference Order | Name |
1 | SHA2-512 |
2 | SHA2-384 |
3 | SHA-1[2] |
Table 10 – Hash algorithm preference order
A hash-based message authentication code (HMAC) provides a value to assess the integrity and authenticity of data. MACs are frequently used when downloading files from the Internet to compare the signatures provided by the host to ensure the integrity of the downloaded files is what was posted by the host.
Preference Order | Name |
1 | AEAD (GCM) |
2 | SHA2-512 |
3 | SHA2-384 |
4 | SHA2-256 |
Table 11 – HMAC preference order
With the above asymmetric and symmetric standards, this has downstream impacts on the configuration of protocols, such as TLS or SSH. In addition, due to the significant change in baseline requirements from TLS version 1.2 to version 1.3, this has proscribed several options which were previously available.
As a result, the following standards are recommended.
Transport Layer Security (TLS) is the successor protocol to Secure Sockets Layer (SSL), which is used to establish secure connections to server-hosted applications. In client/server architecture, servers hold asymmetric keypairs – a public & private key – which is used to identify and trust servers to clients. The most common implementation is for HTTPS certificates for websites.
- TLS 1.3
Due to the changes mentioned above, TLS v1.3 presents a significantly different view in comparison to the TLS v1.2 list. TLS v1.3 cipher suite strings do not specific key exchange methods, as the requirement is to support ‘perfect forward secrecy’ (PFS) which mandates the use of the ECDHE or DHE key exchange methods.
The below list is provided in descending preference order.
Preference Order | Name |
1 | TLS_AES_256_GCM_SHA384 |
2 | TLS_AES_128_GCM_SHA256 |
Table 12 – TLS 1.3 ciphersuite preference order
- TLS 1.2
This table lists the approved TLS v1.2 cipher suites, comprised of a key exchange method, asymmetric algorithm, symmetric algorithm, and hash function, in descending preference order.
Preference Order | Name |
1 | TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384 |
2 | TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA256 |
3 | TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA384 |
4 | TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256 |
5 | TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 |
6 | TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 |
7 | TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 |
8 | TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 |
Table 13 – TLS 1.2 ciphersuite preference order
Secure Shell (SSH) is a command-line interface tool used for interacting and performing administrative tasks on systems, typically Linux. There are open-source and commercially-licensed software that supports this protocol, the most common being OpenSSH. OpenSSH enabled three (3) asymmetric key types out-of-the-box for SSH: ed25519, ECDSA, and RSA.
When implementing SSH, the following sequence of steps should be performed:
- Remove the ed25519 configuration
- Regenerate SSH’s ECDSA private key
- Regenerate SSH’s RSA private key
This will ensure that each SSH private key is randomly generated and avoid re-use of the same private key across multiple systems.
- Secure Shell File Transfer Protocol
Secure Shell File Transfer Protocol (SFTP) is a method of performing file transfer activities over SSH. This is separate to SCP, which also uses SSH for file transfer activities. SFTP configurations should align to the SSH guidance, use separate keys, and block SSH commands from occurring over SFTP connections.
The Internet Key Exchange (IKE) protocol is the foundation to building security associations (SA) for establishing site-to-site VPN connections over Internet Protocol Security (IPsec).
IKE has two active versions, IKEv1 and IKE2.
IKE also has two phases of security association, phase 1 and phase 2. Phase 1 serves to establish the secure communication channel using a key exchange algorithm. Phase 2 then acts to use the secure channel established in phase 1 to negotiate the IPsec communication.
- IKE Phase 1 Minimum Settings
Setting Name | Setting Value |
Operation Mode | IKEv2 |
Authentication | Pre-shared Key |
Encryption | AES-256-CBC[3] |
Hash | SHA-384 |
Diffie-Helman Group | Group 20 |
Lifetime | 14400 seconds |
NAT Traversal | Enabled |
Table 14 – IPSec IKE Phase 1 minimum settings
- IKE Phase 2 Minimum Settings
Setting Name | Setting Value |
Operation Mode | IKEv2 |
Authentication | Pre-shared Key |
IPsec Protocol | ESP |
Encryption | AES-256-GCM |
Hash | SHA-384 |
Diffie-Helman Group | Group 20 |
Lifetime | 14400 seconds |
NAT Traversal | Enabled |
Table 15 – IPSec IKE Phase 2 minimum settings
Asymmetric cryptography is most frequently used for HTTPS certificates and other digital signatures to establish integrity, confidentiality, and authenticity. Certification Authorities (CA), also called Certificate Authorities, are components in the wider Public Key Infrastructure (PKI) used to manage certificates, certificate templates, chains of certification authority trusts, and validation of certificates.
There are two types of Certification Authorities: Root and Intermediate.
Root Certification Authorities are the devices which form the ‘root’ of any chain of certificate hierarchy. The certificates or signatures that get issued to websites always track back up to a Root CA. In any Public Key Infrastructure or certificate chain there is only ever one Root CA for each chain.
Intermediate Certification Authorities are devices which are either issued a digital signature by a Root CA or another Intermediate CA higher in the certificate trust chain, which in turn then also issue certificates.
An Intermediate CA is also known as an ‘issuing CA’, in that opposed to being the root of a trust chain, the Intermediate CA issues out certificates or signatures on behalf of its’ parent Root or Intermediate CA. It is these devices which issue out web server certificates.
When compute is deployed to an environment, whether a virtual machine or container, it will need to receive operating system and application updates to remain secure. Within this environment, there are several different approaches to managing this process.
For the context of this Standard, an “update” is any patch, security or otherwise, which increments the version of the application or operating system. This does not include a version change, such as Windows 10 to Windows 11, or OpenSSL 1.1 to 3.0.
There are several considerations depending on the source operating system for each element.
- Linux operating systems do not have a set update release cycle. Instead, updates are released as soon as they are release-ready.
- Windows operating systems have updates released on a well-known schedule; the second Tuesday of every month (Patch Tuesday). Due to time delays, for Australia, this is a Wednesday.
- Operating system containers are released based on vendor release schedules, which are not necessarily public.
- Application containers are updated in-line with vendor release schedules, which are not necessarily public.
The below individual items will cover standard approaches. Exception management, such as a Linux kernel CVE patch or other equivalent critical fix, will always take effect over these standard approaches.
Virtual Machine images, aka SOE images, will be created once per month.
To enable Windows VM images to receive the latest patches as part of build, the following schedule will be used to create VM images.
Image Version | Schedule |
Windows Client | Friday following Patch Tuesday |
Windows Server | Friday following Patch Tuesday |
Linux (RPM) | First Friday of every month |
Linux (Debian) | First Friday of every month |
Table – Virtual machine image update schedule
Azure Update Management Centre will be used to patch virtual machines. This supports manual and scheduled patching.
The following table describes the proposed standard schedule for updates to VM operating systems. As the virtual machines are not expected to directly face to the Internet, either behind a WAF or otherwise, there is not a sufficient need to patch virtual machines within a few days of a patch release.
Environment | Schedule |
Development | Friday following Patch Tuesday |
Integration | Monday following Development |
UET | Wednesday following Integration |
Training | Wednesday following Integration |
Staging (Disaster Recovery) | Thursday following Training |
Staging (Production) | Wednesday following Disaster Recovery |
Disaster Recovery | Monday following Staging (Disaster Recovery) |
Production | Monday following Staging (Production) |
Table – Virtual machine update schedule
Monday | Tuesday | Wednesday | Thursday | Friday | Saturday | Sunday |
1 | 2 | 3 | 4 | 5 | 6 | 7 |
8 | 9 | 10 | 11 | 12 | 13 | 14 |
15 | 16 | 17 | 18 | 19 | 20 | 21 |
22 | 23 | 24 | 25 | 26 | 27 | 28 |
29 | 30 | 31 |
- 12th of the month: Development
- 15th of the month: Integration
- 17th of the month: UET + Training
- 18th of the month: Staging (Disaster Recovery)
- 22nd of the month: Disaster Recovery
- 24th of the month: Staging (Production)
- 29th of the month: Production
Container image updates are pushed or released on a schedule determined by the developer or vendor of the specific container.
Platform container images will be re-tagged and pushed to a hosted container registry within 2 business days of release.
Network Detailed Design
Azure Cloud Platform |
Contents
1 Target State Network Design. 6
1.1 Hub-and-spoke network topology. 6
1.2 IP address supernet allocation. 7
1.5.2 Virtual Network Gateway route table. 11
1.5.3 Kubernetes subnets route table. 11
1.6 Network security groups. 12
1.6.2 Jira application subnets. 13
1.8 IP address management tooling. 13
1.9 ExpressRoute Connectivity. 13
1.9.1 ExpressRoute Circuit. 14
1.9.2 Virtual Network Gateway. 14
1.9.3 Peering Configuration. 14
1.9.4 Palo Alto VM-Firewall 15
2.1 DNS Zone Configuration. 15
2.1.2 Internal DNS Zone(s). 16
2.1.4 Subdomain delegation. 16
3 Monitoring Customer Experience. 18
Tables
Table 2 – CBX Contact Details. 2
Table 3 – CBX Version Control Details. 2
Table 4 – CBX Version Control Changelog. 2
Table 5 – CIDR Supernet allocation. 7
Table 6 – CIDR allocation breakdown per environment tier. 7
Table 7 – Resource groups for Production network elements. 7
Table 8 – Resource groups for Disaster Recovery network elements. 8
Table 9 – Resource groups for Development network elements. 8
Table 10 – Resource groups for Sandpit network elements. 8
Table 11 – Virtual Networks and CIDR allocations. 8
Table 12 – Hub and spoke network peering assignments. 9
Table 13 – Connectivity Virtual Network subnets. 9
Table 14 – Identity Virtual Network subnets. 9
Table 15 – Security Virtual Network subnets. 9
Table 16 – Integration Virtual Network subnets. 10
Table 17 – Management Virtual Network subnets. 10
Table 18 – Atlas Virtual Network subnets. 10
Table 19 – Default Route Table configuration. 10
Table 20 – Example Default Route Table configuration. 11
Table 21 – VNG Route Table configuration. 11
Table 22 – K8s Route Table configuration. 12
Table 23 – Example K8s Route Table configuration. 12
Table 24 – AVD non-default NSG rules. 13
Table 25 – Active Public IP prefixes. 13
Table 26 ExpressRoute Circuit. 14
Table 27 Virtual Network Gateways. 14
Table 28 ExpressRoute Pre-requisites. 15
Table 29 Palo Alto VM Appliances. 15
Table 30 Environment domains in use. 16
Table 31 Internal DNS Domain. 16
Table 32 External DNS domains per environment. 16
Figures
Figure 1 – Hub-and-spoke Topology. 6
18 Target State Network Design
This section describes the target, or end-state, design and topology of the network within the Azure environment.
18.1 Hub-and-spoke network topology
The below diagram illustrates the design intent for the hub-and-spoke network for this environment, and for the environment hubs for each environment tier.
Figure 1 – Hub-and-spoke Topology
18.2 IP address supernet allocation
The below table identifies the total IP address range allocation for Azure in Classless Inter-Domain Routing (CIDR) format. The CIDR format is used throughout this document.
Azure allocation type | CIDR allocation |
All regions | 10.216.0.0/13 |
Table 5 – CIDR Supernet allocation
18.2.1 Allocation breakdown
Note that not all the environments below will be used. They have been provided in advance for network allocation purposes only.
Environment allocation | Azure region | CIDR allocation |
Production | Australia East | 10.216.0.0/16 |
Disaster Recovery | Australia Southeast | 10.217.0.0/16 |
Non-Production | Australia East | 10.218.0.0/16 |
Sandpit | Australia East | 10.219.0.0/16 |
Table 6 – CIDR allocation breakdown per environment tier
18.3 Resource groups
The below sections identify all the resource groups in each Azure subscription where network elements will be deployed. Detail into each subscription name and type can be found in the Tenant Hierarchy Solution Design document.
18.3.1 Production
Resource Group name | Function | Subscription |
rg-CUST-con-prd-ae-eslz-001 | ESLZ – Network Hub | CUSTPD-Connectivity |
rg-CUST-idm-prd-ae-eslz-001 | ESLZ – Identity resources | CUSTPD-Identity |
rg-CUST-mgm-prd-ae-eslz-001 | ESLZ – Management resources | CUSTPD-Management |
rg-CUST-sec-prd-ae-eslz-001 | ESLZ – Security resources | CUSTPD-Security |
rg-CUST-int-prd-ae-eslz-001 | ESLZ – Integration resources | CUSTPD-Integration |
rg-CUST-jira-pd-ae-app-001 | Application – Jira Resources | CUSTPD-Jira |
Table 7 – Resource groups for Production network elements
18.3.2 Disaster Recovery
Resource Group name | Function | Subscription |
rg-CUST-con-drp-as-eslz-001 | ESLZ – Network Hub | CUSTPD-Connectivity |
rg-CUST-mgm-drp-as-eslz-001 | ESLZ – Management resources | CUSTPD-Management |
rg-CUST-int-drp-as-eslz-001 | ESLZ – Integration resources | CUSTPD-Integration |
rg-CUST-jira-drp-as-app-001 | Application – Jira Resources | CUSTPD-Jira |
Table 8 – Resource groups for Disaster Recovery network elements
18.3.3 Non-Production
Resource Group name | Function | Subscription |
rg-CUST-con-npd-ae-eslz-001 | ESLZ – Network Hub | CUSTNP-Connectivity |
rg-CUST-mgm-npd-ae-eslz-001 | ESLZ – Management resources | CUSTNP-Management |
rg-CUST-int-npd-ae-eslz-001 | ESLZ – Integration resources | CUSTNP-Integration |
rg-CUST-jira-npd-ae-app-001 | Application – Jira resources | CUSTNP-Jira |
Table 9 – Resource groups for Development network elements
18.3.4 Sandpit
Resource Group name | Function | Subscription |
rg-CUST-sdp-npd-ae-eslz-001 | ESLZ – Sandpit Resources | CUSTNP-Sandpit |
Table 10 – Resource groups for Sandpit network elements
18.4 Virtual networks
Virtual Network name | Subscription | CIDR allocation |
vn-con-ae-prd-eslz-001 | CUSTPD-Connectivity | 10.216.0.0/21 |
vn-idm-ae-prd-eslz-001 | CUSTPD-Identity | 10.216.8.0/21 |
vn-mgm-ae-prd-eslz-001 | CUSTPD-Management | 10.216.16.0/21 |
vn-sec-ae-prd-eslz-001 | CUSTPD-Security | 10.216.24.0/21 |
vn-int-ae-prd-eslz-001 | CUSTPD-Integration | 10.216.32.0/21 |
vn-app-ae-prd-jira-001 | CUSTPD-Jira | 10.216.40.0/21 |
vn-con-as-drp-eslz-001 | CUSTPD-Connectivity | 10.217.0.0/21 |
vn-mgm-as-drp-eslz-001 | CUSTPD-Management | 10.217.16.0/21 |
vn-sec-as-drp-eslz-001 | CUSTPD-Security | 10.217.24.0/21 |
vn-int-as-drp-eslz-001 | CUSTPD-Integration | 10.217.32.0/21 |
vn-app-ae-drp-jira-001 | CUSTPD-Jira | 10.217.40.0/21 |
vn-con-ae-npd-eslz-001 | CUSTNP-Connectivity | 10.218.0.0/21 |
vn-mgm-ae-npd-eslz-001 | CUSTNP-Management | 10.218.16.0/21 |
vn-int-ae-npd-eslz-001 | CUSTNP-Integration | 10.218.32.0/21 |
vn-jira-ae-np-eslz-001 | CUSTNP-Jira | 10.218.40.0/21 |
vn-sdp-ae-eslz-001 | CUSTNP-Sandpit | 10.219.0.0/21 |
Table 11 – Virtual Networks and CIDR allocations
18.4.1 Peering
Hub network | Spoke network(s) | ||||||
vn-con-ae-prd-eslz-001 |
| ||||||
vn-con-as-drp-eslz-001 |
| ||||||
vn-con-ae-np-eslz-001 |
|
Table 12 – Hub and spoke network peering assignments
18.4.2 Subscription subnets
Each CIDR allocation in the below sections and tables only identify the third (3rd) and fourth (4th) octet. This is due to the first (1st) and second (2nd) octets being consistent across each environment. The first and second octets are identified as <CIDR Prefix> in each row.
18.4.2.1 Connectivity
Subnet Name | Route Table | NSG | Function | CIDR allocation |
AzureFirewallSubnet | No | No | Azure Firewall | <CIDR Prefix>.2.0/23 |
AzureFirewall ManagementSubnet | No | No | Azure Firewall | <CIDR Prefix>.4.0/24 |
GatewaySubnet | Yes | No | Virtual Network Gateways | <CIDR Prefix>.5.0/24 |
Table 13 – Connectivity Virtual Network subnets
18.4.2.2 Identity
Subnet Name | Route Table | NSG | Function | CIDR allocation |
sn-default | Yes | Yes | n/a – deployment only | <CIDR Prefix>.8.0/29 |
Table 14 – Identity Virtual Network subnets
18.4.2.3 Security
Subnet Name | Route Table | NSG | Function | CIDR allocation |
sn-default | Yes | Yes | n/a – deployment only | <CIDR Prefix>.24.0/29 |
Table 15 – Security Virtual Network subnets
18.4.2.4 Integration
Subnet Name | Route Table | NSG | Function | CIDR allocation |
sn-default | Yes | Yes | n/a – deployment only | <CIDR Prefix>.32.0/29 |
Table 16 – Integration Virtual Network subnets
18.4.2.5 Management
Subnet Name | Route Table | NSG | Function | CIDR allocation |
sn-avd | Yes | Yes | Azure Virtual Desktop | <CIDR Prefix>.16.0/26 |
sn-aks-sys | No | No | AKS system nodes | <CIDR Prefix>.17.0/24 |
sn-aks-usr | No | No | AKS workload hosting nodes | <CIDR Prefix>.18.0/23 |
sn-aks-lb | No | No | AKS internal load balancer | <CIDR Prefix>.20.0/26 |
sn-packer | Yes | Yes | Packer VM builds | <CIDR Prefix>.20.64/29 |
sn-images | Yes | Yes | ACR private endpoint | <CIDR Prefix>.20.72/29 |
sn-ado-agents | Yes | Yes | Azure DevOps agents | <CIDR Prefix>.20.80/29 |
Table 17 – Management Virtual Network subnets
18.4.2.6 Jira
Subnet Name | Route Table | NSG | Function | CIDR allocation |
sn-default | Yes | Yes | n/a – deployment only | <CIDR Prefix>.40.0/29 |
Table 18 – Atlas Virtual Network subnets
18.5 Route tables
As part of the hub-and-spoke implementation, subnets in Virtual Networks need to be instructed where to send traffic. This is achieved in this design through User-Defined Routes (UDRs), implemented through Route Tables.
The below table defines the default route table attached to each subnet except where noted in subsequent sections.
Route Table property | Value | ||||||||||||
Routes |
| ||||||||||||
Propagate Gateway Routes | Disabled |
Table 19 – Default Route Table configuration
The Vnet-Local route is required to be added to ensure that inter-subnet traffic is inspected by the firewall. Without this route, traffic in one subnet can pass directly to another subnet in the same virtual network.
18.5.1.1 Example Route Table for Nonproduction Management Subscription
Route Table property | Value | ||||||||||||
Routes |
| ||||||||||||
Propagate Gateway Routes | Disabled |
Table 20 – Example Default Route Table configuration
18.5.2 Virtual Network Gateway route table
This route table defines the route table attached to the ‘GatewaySubnet’ in the Connectivity hub virtual network.
This specific table requires Internet access directly, as the Virtual Network Gateways require Public IP addresses, and, that traffic for each Virtual Network CIDR traverses the local region Azure Firewall.
Route Table property | Value | ||||||||||||||||||||||||
Routes |
| ||||||||||||||||||||||||
Propagate Gateway Routes | Enabled |
Table 21 – VNG Route Table configuration
18.5.3 Kubernetes subnets route table
Due to the architecture of separating AKS system nodes from user nodes, to ensure separation of workloads and optimise network utilisation for scaling, Kubernetes route tables should be configured to ensure individual subnets communicate directly rather than traversing an inspection appliance.
This is particularly important for autoscaling nodes, where the IP address of each node may change over small intervals of time.
Route Table property | Value | ||||||||||||||||||||||||
Routes |
| ||||||||||||||||||||||||
Propagate Gateway Routes | Disabled |
Table 22 – K8s Route Table configuration
18.5.3.1 Example Route Table for Nonproduction Management Kubernetes subnets
Route Table property | Value | ||||||||||||||||||||||||
Routes |
| ||||||||||||||||||||||||
Propagate Gateway Routes | Disabled |
Table 23 – Example K8s Route Table configuration
18.6 Network security groups
As a rule, each subnet requires a Network Security Group attached to serve as a Layer 4 access control list. Given the use of a firewall as the traffic inspection hub, NSGs serve as the “last resort” option to limit inbound traffic.
Some subnets, however, do not or should not use NSGs. This includes subnets such as the Azure Firewall subnet, as that is a dedicated subnet for the Firewall which should be receiving all traffic.
NSGs also can leverage Application Security Groups (ASGs) for rules, in addition to IP addresses and Azure Service Tags. ASGs are best used attached to VM interfaces, rather than IP addresses.
18.6.1 ESLZ subnets
NSGs deployed to ESLZ subnets will implement the standard default rules as provided in Network Security Groups.
For services exposed to the Internet through Destination Network Address Translation (DNAT) rules on Azure Firewall, custom rules will be implemented both in Azure Firewall and on the specific NSG. For example, SFTP file transfer.
18.6.2 Jira application subnets
The Jira application will be deployed into a secure environment, leveraging NSG’s that only allow the minimum network traffic required for the environment to operate. The below are examples of NSG’s that will be deployed to provide granular network controls over ingress/egress traffic to each subnet.
18.6.2.1 Jira AVD subnet
Name | Source | Destination | Priority # |
In-FromEslzAvd | <ESLZ AVD subnet> | VirtualNetwork | 150 |
In-DefaultDeny | Any | any/* | 3999 |
Table 24 – AVD non-default NSG rules
18.7 Public IP prefixes
Public IP prefixes are procured to be able to define a contiguous block of IP addresses from within Azure. This allows for simpler IP-based rulesets, rather than managing single IP address entries.
The following IP Prefixes will be used by the VNET gateway to establish an external connection point over ExpressRoute to the CUST Core network.
Public IP Prefix name | Environment | CIDR range |
pfx-con-ae-prd-eslz-001 | Production | x.x.x.x/30 |
pfx-con-as-drp-eslz-001 | Disaster Recovery | x.x.x.x/30 |
pfx-con-ae-dev-eslz-001 | Nonproduction | x.x.x.x/30 |
Table 25 – Active Public IP prefixes
The CIDR ranges will be deployed using Microsoft provided Public IP’s. The actual IP range will be defined in the As-Built documentation once they have been deployed.
18.8 IP address management tooling
To be able to keep track of IP addressing, both public and private, as well as network assets, a suite of tools known as ‘IP Address Management’ tools are used to track and manage allocations. This environment will implement leverage the existing CUST IPAM tooling solution, phpIPAM.
Deployment details for IPAM will be captured in the As-Build documentation.
18.9 ExpressRoute Connectivity
ExpressRoute connectivity will be established using Megaport as the supported connectivity provider, this will ensure traffic traverses dedicated private links and not the internet.
18.9.1 ExpressRoute Circuit
An ExpressRoute Circuit consists of two connections to two Microsoft Enterprise Edge Routers (MSEE’s) at an ExpressRoute location from the edge network. CUST will leverage the ExpressRoute partner model, meaning that connectivity will be provided by the approved ExpressRoute partner.
A single ExpressRoute circuit will be deployed with the below configuration. Microsoft recommends deploying minimum 2 ExpressRoute Circuits to support high availability, however given the size and criticality of the environment in current state, this will not be deployed.
Circuit | Deployment Detail |
Subscription | CUSTPD-Management |
Resource Group | vn-con-ae-prd-eslz-001 |
Region | Australia East |
Name | erc-con-ae-prd-eslz-001 |
18.9.1.1 Service Key
Once the ExpressRoute circuit has been successfully created, an S-Key (Service Key) will be issued to the circuit by Microsoft automatically. The connectivity provider will use the S-Key to complete the provisioning process between their network and Microsoft Azure.
18.9.2 Virtual Network Gateway
A virtual network gateway serves two purposes: exchange IP routes between the networks and route network traffic. The connectivity hub in each logical environment (PRD, DR, NPD) will have a dedicated subnet, called ‘GatewaySubnet’, deployed to support the virtual network gateway. The naming convention applied here lets Azure know to deploy VNet Gateway VM’s and services directly into the subnet.
The following details capture each virtual network gateway that will be deployed.
Detail | Prod Gateway | DR Gateway | Non-Prod Gateway |
Subscription | CUSTPD-Management | CUSTPD-Management | CUSTNP-Management |
Connected Hub | vn-con-ae-prd-eslz-001 | vn-con-as-drp-eslz-001 | vn-con-ae-npd-eslz-001 |
Name | vng-con-ae-prd-eslz-001 | vng-con-as-drp-eslz-001 | vng-con-ae-npd-eslz-001 |
Table 27 Virtual Network Gateways
18.9.3 Peering Configuration
Private peering will establish a secure connection to the CUST core network from Azure, allowing bi-directional connectivity between networks. More detail will be provided in the As-Built documentation however the following items are pre-requisites which are required for implementation of ExpressRoute circuits.
The following detail will need be captured prior to the setup of ExpressRoute.
Pre-Requisite | Detail |
1 x /29 Subnet OR 2x /30 Subnets | Required to setup the primary and secondary routing interfaces. These IP CIDR ranges must not overlap with the Azure ranges in use (RFC1918). |
VLAN ID | A single VLAN ID will be used by both Primary and secondary links. This is used to establish peering with the ExpressRoute circuit. |
AS Number | Unique IANA assigned number used to establish peering with ExpressRoute. |
Core Network Edge routes advertised. | Existing network edge routers must have their routes advertised to Azure via BGP when peering is being configured. |
Table 28 ExpressRoute Pre-requisites
18.9.4 Palo Alto VM-Firewall
CUST will extend the use of its strategic Palo Alto Firewall security services into Azure. Palo Alto Next-Generation Firewalls will be leveraged for VPN connectivity into the DR landing zone due to limitation with Virtual Network Gateways in the Australia Southeast region, and a Palo NGFW will be deployed for temporary internet access for the Azure landing zones until on-premises networks are redeployed to support the landing zone connectivity requirements.
Virtual Machine | Configuration | Detail |
Panorama VM | panconaeprd001 (primary) panconasdrp001 (secondary) | Vnets:vn-con-ae-prd-eslz-001, vn-con-as-drp-eslz-001 Image Name: Panorama (BYOL) – x64 Gen1 Size: Standard_D16s_v5 |
Palo Alto NGFW VM | pangfwintaeprd001 (temporary internet gateway) pangfwconasdrp001 (DR gateway) | Vnets:vn-internet-ae-prd-eslz-001, vn-con-as-drp-eslz-001 Image Name: VM Series Flex Bundle 2 Size: Standard DS3_v2 |
Table 29 Palo Alto VM Appliances
Configuration details for the Palo Alto VM appliances will be detailed in the As-Built documentation.
19 Name Services
19.1 DNS Zone Configuration
This section describes the configuration of internal and external DNZ zones, where they are configured and how environment isolation is achieved.
19.1.1 Domain(s)
The domains used for DNS for this environment are described in the following table. These domains may be top-level domains, or sub-domains delegated to this environment.
DNS Domain(s) | Domain Type |
CUST.com.au | Top level domain |
azure.CUST.com.au | Subdomain |
Table 30 Environment domains in use.
19.1.2 Internal DNS Zone(s)
The following table lists the DNS zones used for internal namespaces within the environment.
DNS Domain(s) | AD Tenant |
azure.CUST.com.au | CUST.com.au |
These DNS zones will be made available over the internal network, for the named environment, at the indicated IP address(es).
Internal DNS Domain | DNS Environment | DNS Server IP Address(es) |
sandpit.azure.CUST.com.au | Sandpit | Firewall DNS Proxy IP |
np.azure.CUST.com.au | Non-Production | Firewall DNS Proxy IP |
azure.CUST.com.au | Production / Disaster Recovery | Firewall DNS Proxy IP |
Table 34 – Internal DNS domain per environment
19.1.3 Public DNS Zone(s)
Public DNS domains will not be configured as part of this engagement as this ESLZ will not be accessible from the internet.
19.1.4 Subdomain delegation
Each subscription or environment within Azure is intended to have its own subdomain, to isolate that environment from production. The following table describes the breakdown of each subdomain and the subscription that subdomain zone is delegated to.
Public DNS Subdomain | Environment | Subscription Name |
sandpit.azure.CUST.com.au | Sandpit | CUSTNP-Sandpit |
dev.azure.CUST.com.au | Non-Production (Shared Services) | CUSTPD-Connectivity |
prd.azure.CUST.au dr.azure.CUST.com.au | Production / Disaster Recovery (Shared Services) | CUSTPD-Connectivity |
Table 32 External DNS domains per environment
19.2 Private endpoints
In addition to “normal” DNS resolution functions, Azure has a specific capability around the use of private network connections for Platform-as-a-Service resources known as ‘private endpoints.’ The use of private endpoints requires an integrated, specially named DNS zone for each different resource type.
Inbound and Outbound endpoints will be configured to assist in creating an end-end hybrid DNS architecture, allowing for private DNS namespaces to be resolved to the extended CUST core network. The detail of these endpoints will be documented in the As-Built documentation.
Governance and Mgt As Built
Azure Cloud Platform |
Contents
3 Governance and compliance. 6
3.1.1 Azure policies and Initiatives. 6
3.2 Data classification and tagging. 6
3.3 End-user access control mechanisms. 7
4 Monitoring Customer Experience. 8
Tables
Table 2 – RLX Contact Details. 2
Table 3 – RLX Version Control Details. 2
Table 6 – Azure policies and Initiatives. 6
Table 12 – End-user access control mechanisms. 7
Table 13 – Log data retention. 7
20 Summary
20.1 Background
The Australian CUST Corporation (CUST) has engaged RLX to establish an Azure Enterprise-Scale Landing Zone to provide a foundational architecture for proof-of-concept testing. This project serves as the first steppingstone, kick starting CUST’s feasibility studies into their cloud adoption journey. The Azure platform will act as the CUST strategic platform for short-term proof-of-concept testing, with the intent for CUST to conduct feasibility studies and determine a path forward in transitioning existing applications from on-prem to the cloud. Further details pertaining to this project can be referenced in SoW— 2023050957129.
20.2 Purpose
The purpose of this document is to document the build and configuration properties of CUST’s governance policies applied in the Azure Enterprise Scale Landing Zone (ESLZ). It is expected that those reading this document will have an intermediate level of technical understanding of Azure and associated components that contribute towards the overall ESLZ.
20.3 In-scope
In scope |
ABAC-Governance and Management |
Leverage existing DTD’s including:
|
21 Requirements
21.1 Technical Requirements
As per the High-Level Solution Design, the Governance solution design should inherit the following requirements.
ID | Design Decision | Rationale |
DD-ELZ-20 | Separation of Duties | Separate privileged and non-privileged accounts will be used to isolate. |
DD-ELZ-53 | Compliance Framework(s) Alignment | Configure resources and services to comply with one or more compliance frameworks. |
DD-ELZ-54 | Cloud Security Posture Management Tooling | Microsoft Defender will be deployed as the CSPM solution for the environment. |
CR-1 | Platform-wide classification | The platform, including the ESLZ, will be built to known-good patterns, which have been designed to meet a stringent controls framework such as the ISM. |
CR-2 | Data sovereignty | The platform, including the ESLZ, is required to be hosted within Australian borders. |
CR-3 | Regulatory compliance | The platform, including the ESLZ, is required to meet the following regulatory framework(s): · IRAP PROTECTED |
CR-4 | Security compliance | The platform, including the ESLZ, is required to align to the following security compliance framework(s):
· IRAP PROTECTED · Secure Cloud Computing Architecture (SCCA) |
22 Governance and compliance
22.1 Enforcement mechanisms
22.1.1 Azure policies and Initiatives
The table below details the Azure policies that have been deployed into the CUST environment. This is broken down into where the policy sits at the parent level.
Parent | Policy/Initative name | Function | Description |
Mg-root (Management Group) | Mandatory Tagging Policy | DeployIfNotExists | Enforces mandatory tags on all resources across the ESLZ. |
Australian ISM Protected | Audit | This policy audits all resources and assesses them against the AUS ISM Protected. | |
Allowed Locations | Deny | Deny Locations outside of Australia East or Australia SouthEast for resources. | |
Stream Az_Act_logs_LAW | DeployIfNotExists | This policy enforces all Azure Activity Logs be streamed to the sentinel log analytics workspace. | |
Subscription level (all existing subscriptions) | CUST Default | Audit | Audit resources against Azure Security Center best practice polices. |
Table 6 – Azure policies and Initiatives
22.2 Data classification and tagging
22.2.1 Resource Tagging
Tagging standards, as stipulated by CUST, have been enforced via Azure policy. At present, resources will inherit tags at a subscription level, with room for this to expand down to the resource group level. The current policy in place enforces tag inheritance. The table below details which tags will be inherited by resources by default based on this policy.
Tag name | Description |
Costcenter | This tag identifies the costcenter in which the resource is being billed to. |
ServiceOwner | This tag identifies the service owner that is responsible for managing the overall service. |
BusinessUnit | This tag identifies the business unit in which the resource is managed by. |
ResourceOwner | This tag identifies the owner for the specific resource that has been deployed. Note* this differs from the service owner. |
22.2.2 Data classification
As per CR-4, the CUST Azure environment has been designed and engineered to only store data to the level of Protected. The onus will be placed upon CUST staff and authorized users to classify data appropriately.
22.3 End-user access control mechanisms
The table below details all end-user access control mechanisms that have been deployed and configured.
Mechanism |
Conditional Access Policies |
Role-based Access |
Entra ID Privileged Identity Management |
Table 12 – End-user access control mechanisms
22.4 Log data retention
The table below details all data retention properties, where applicable, for audit and compliance logs collected. CommVault was discussed as the potential long term storage solution for cloud data, however this has not been implemented in the scope of this project.
Log type | Retention time |
Diagnostic logs | Interactive retention: 30 days Archive period : 7 years |
Security logs | Interactive retention: 30 days Archive period : 7 years |
[1] Source: Kubernetes Conceptual Overview, Kubernetes.io, https://kubernetes.io/docs/concepts/overview/
[2] Note that SHA-1 is listed here as there are several cases where SHA-1 will remain, such as IKE Phase 1 security association hashing and thumbprinting of digital signatures or X.509v3 certificates. All other uses are prohibited.
[3] [3] Currently, CBC is the only cipher available for IKE Phase 1 authentication. This is the only exception to the approved cipher list.
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.