Cloud and Infrastructure FAQ: Architecture, Ownership, Security and Cost

A cloud and infrastructure FAQ explaining service models, shared responsibility, reliability, security, migration, cost controls and provider accountability.

Edilec Research Updated 2026-07-14 Cloud & DevOps

A cloud and infrastructure FAQ should help leaders decide what service they need, which responsibilities remain theirs and what evidence demonstrates that the platform is reliable, secure and economically controlled. Cloud infrastructure includes accounts or subscriptions, identity, networks, compute, storage, data services, observability, backup, recovery and the operating practices around them. The business outcome is not possession of cloud resources; it is a supportable service whose user journeys survive expected change and recover from failure.

This guide complements Edilec's cloud infrastructure delivery plan, implementation checklist and infrastructure services FAQ. Cloud provider capabilities and prices change, so validate product-specific decisions against current documentation. The principles here focus on architecture and ownership that remain important across providers.

What does cloud infrastructure include?

Scope the service from user-facing workloads inward. List applications, data, integrations, locations, critical periods, recovery needs and compliance constraints. Then identify landing zones, account structure, identity, network connectivity, edge protection, compute platforms, storage, databases, keys, secrets, logging, backup and support. Include DNS, domains, certificates, source and build systems where an outage would block operation or recovery. A proposal that ends at virtual machines or Kubernetes has omitted the operating service.

The service model may be infrastructure as a service, managed platform, software as a service or a combination. More managed capability can reduce undifferentiated operational work, but it changes configuration, portability, limits and supplier dependency rather than eliminating responsibility. Compare options using workload needs, team capability, recovery, security, data gravity, latency, commercial terms and exit. The AWS, Azure and Google well-architected frameworks all organize design around multiple quality attributes rather than a single technology preference.

Service areaCustomer decisionEvidence
Business boundaryWhich journeys, data and hours matter?Service definition and dependency map
FoundationHow are identity, accounts and networks governed?Versioned landing-zone controls
ProtectionWho owns access, keys, patching and response?Responsibility and control matrix
ReliabilityWhat fails, recovers and degrades?SLOs, restore and failover exercises
EconomicsWho allocates, forecasts and optimizes spend?Unit cost and variance review
ExitHow are data, domains and operation transferred?Tested export and transition plan

Who is responsible for security and operations?

Responsibility is shared but not vague. Providers secure particular facilities and managed-service layers; customers still configure identity, networks, data access, logging, resilience and workloads within the contracted model. Managed service partners may operate selected controls, but the business remains accountable for risk acceptance. Create a control-level matrix naming provider, partner and customer tasks, evidence, frequency and escalation. Include subcontractors and support access. Review the matrix when architecture or service terms change.

Cloud service layers
A dependable cloud service connects business need, governed foundations, protection, reliable operation, economics and improvement.

Use a common outcomes vocabulary such as the NIST Cybersecurity Framework 2.0, which includes Govern, Identify, Protect, Detect, Respond and Recover. It helps expose proposals that emphasize prevention but omit recovery or governance. Retain organizational ownership of root identity, domains, billing, critical keys, policy and evidence access. A provider should not be the sole holder of assets needed to replace that provider.

How should cloud reliability be designed?

Start with service level indicators for successful customer journeys, latency, correctness, freshness and durability. Define objectives and tolerated impact, then design redundancy and recovery that meet them. High availability across zones does not automatically survive region, identity, deployment or data-corruption failure. Map dependency failure modes and choose graceful degradation. Limit the blast radius with accounts, cells, queues, bulkheads, quotas and progressive release. Keep capacity for failover, not only ordinary peak.

Backups are useful only when restorable. Define recovery time and recovery point by data set, protect copies from the same identities and failure domain, and exercise restore into an isolated environment. Reconcile application consistency, not only object count. Test loss of connectivity, credentials, dependency and zone; include the people and communications needed during an incident. The Azure Well-Architected Framework and AWS Well-Architected Framework provide provider-specific review guidance, but drills supply local proof.

Which cloud security controls matter first?

Establish centralized identity, phishing-resistant authentication for privileged users where appropriate, least privilege, separate administrative roles and emergency access. Use policy-as-code or equivalent guardrails for region, network exposure, encryption, logging and approved services. Protect secrets outside source code and rotate them. Inventory assets and data, scan for exposure and vulnerabilities, and prioritize remediation by reachable consequence. Keep development and production authority separate while preserving an auditable emergency path.

Collect control-plane, identity, network and workload evidence into a monitored path. OpenTelemetry signals provide a vendor-neutral basis for application traces, metrics and logs, complementing native cloud audit sources. Define who responds to each alert and test the path. Avoid collecting sensitive payloads by default. Security posture dashboards are hypotheses about configuration; validate them with access reviews, attack-path analysis, recovery exercises and incident learning.

How should workloads move to cloud infrastructure?

Classify each workload to retain, retire, replace, rehost, replatform or refactor based on business value and constraints. Build the governed foundation before migration. Discover dependencies, data flows, licenses, support windows and operational owners. Pilot a representative but bounded workload, not the easiest irrelevant system. Rehearse data transfer, reconciliation, cutover, communication and rollback. A migration is complete only when the cloud service is supportable and old risk, access and cost are retired.

Use waves based on dependency and business risk. Set entry criteria for architecture, security, observability, recovery, capacity, support and customer readiness. Maintain temporary connectivity and dual operation deliberately; unmanaged hybrid periods create permanent cost and attack surface. After each wave, measure incidents, performance, cost, support demand and recovery evidence. Correct the factory before increasing throughput. Preserve data and configuration portability where exit consequence justifies it.

How are cloud costs controlled without harming service?

Allocate spend with account structure and consistent metadata, then connect it to products, environments and units such as active customer, transaction or stored terabyte. The FinOps Framework describes a collaborative practice joining engineering, finance and business stakeholders. Forecast demand, investigate variance and assign anomaly response. Optimize idle resources, data lifecycle, architecture, licenses and purchase commitments only after understanding service and resilience effects. A lower bill paired with slower recovery is not necessarily efficiency.

QuestionWeak answerDecision-grade evidence
Is the service reliable?Provider uptime percentageJourney SLOs, incidents and restore results
Is access controlled?Roles are configuredEntitlement review and denied-path tests
Is spend efficient?Monthly bill decreasedUnit cost with demand and service context
Can the provider be changed?Data can be exportedTimed export, dependencies and transition rights
Is monitoring effective?Many alerts existActionable detection and exercised response
Is migration complete?Traffic movedLegacy cost, access and risk are retired

Evaluate providers and partners through workload evidence. Confirm service locations, limits, support tiers, incident communication, subcontractors, data use, audit material and roadmap dependencies. Test support with a representative case before a crisis. Contracts should align service definitions, security duties, notification, data return, deletion, price-change mechanisms, termination and transition assistance. Certifications can support due diligence but do not prove that the customer's architecture and configuration are controlled.

Design exit while choices remain. Inventory proprietary services and quantify replacement effort. Keep application data in documented schemas, automate environment definitions and retain access to logs and configuration history required for transition. Periodically test export completeness and restoration outside the primary managed service for the most consequential data. Exit readiness does not require lowest-common-denominator architecture; it requires an informed decision about dependency, value and recoverability.

Build internal capability around service ownership, architecture, identity, security, reliability, cost and supplier management. A small organization may use a managed provider for round-the-clock operation, but it still needs qualified people to challenge evidence and make risk decisions. Define on-call and escalation across customer and provider teams. Exercise handoffs so incidents do not stall at contract boundaries. Preserve runbooks and access when key individuals leave.

Review the cloud service quarterly or after material change. Examine user objectives, incidents, restore evidence, security exposure, capacity, unit cost, provider changes and improvement work. Revisit assumptions about region, service tier and portability. Prioritize recurring failure and manual toil rather than adding more dashboards. The review should fund corrective action and retire obsolete controls; governance without a route to change becomes reporting overhead.

Keep architecture decisions linked to workload facts. Record the option chosen, alternatives, service limits, security and reliability assumptions, cost model and revisit trigger. Reassess the records with service owners after major demand change, provider deprecation or incident. This prevents fashionable patterns and inherited defaults from becoming permanent. It also gives auditors, operators and future architects a concise explanation of why the service is built as it is.

Key takeaways

  • Define cloud infrastructure as an operated business service, not a resource list.
  • Assign shared responsibilities at control and evidence level.
  • Design reliability from journey objectives and exercise restore and failover.
  • Connect identity, policy, telemetry and incident response across the stack.
  • Measure unit economics and retire legacy cost and risk after migration.

Frequently asked questions

Is cloud always cheaper? No; compare total operating cost and value under realistic demand. Does multi-cloud remove lock-in? It may reduce selected dependencies while adding complexity; target portability where consequence warrants it. Are managed services more secure? They can reduce some operational burden, but customer configuration and data responsibilities remain. How often should recovery be tested? On a risk-based schedule and after material change, using representative data and people. What should remain customer-owned? At minimum, governance, risk acceptance and controlled access to identity, domains, billing, data and evidence needed to operate or exit.

Conclusion

Cloud infrastructure becomes valuable when architecture, ownership and operation form one service. Clear responsibility, journey-based reliability, tested security and recovery, disciplined migration and visible unit economics let teams use provider capabilities without surrendering control. The strongest evidence is operational: the service can change, fail, recover and exit in ways the organization has actually rehearsed.

Continue with related articles