An infrastructure services FAQ should help a buyer determine what must be operated, who remains accountable and how the service will behave when change or failure arrives. Infrastructure services include the people, processes and technical foundations that keep compute, storage, networks, identity, observability, backup and recovery available to the applications a business depends on. The category may cover cloud, on-premises systems or a hybrid estate. The useful buying unit is not a collection of servers; it is a defined service with measurable outcomes, explicit boundaries and a support model.
This guide answers commercial and technical questions for teams comparing internal delivery, a managed provider or a shared model. It complements Edilec's infrastructure services delivery plan, implementation checklist and managed infrastructure guide. Use those guides after the questions here establish the service boundary and decision criteria.
What do infrastructure services include?
A complete scope normally covers asset and configuration records, account and network foundations, identity administration, operating-system and platform maintenance, capacity, monitoring, backup, recovery, vulnerability response, change control, incident response and supplier coordination. Application support may be included, but it should never be assumed. Ask whether the provider owns only infrastructure health or also traces user-facing failures through databases, queues, APIs and third-party dependencies. The answer determines whether an incident is resolved end to end or passed between teams.

The NIST Cybersecurity Framework 2.0 offers a useful outcome vocabulary across Govern, Identify, Protect, Detect, Respond and Recover. It is not a product specification, but it exposes gaps in a proposal that emphasizes provisioning while omitting governance, detection or recovery. Tailor the required outcomes to the service's risk and business importance; a public checkout platform needs a different control depth and support window from a temporary internal development environment.
| Service area | Questions to settle | Acceptance evidence |
|---|---|---|
| Business boundary | Which journeys, locations and hours are covered? | Approved service definition and dependency map |
| Platform | Which cloud accounts, networks, hosts and data services are managed? | Current inventory and configuration baseline |
| Protection | Who owns access, patching, secrets, backup and exceptions? | Control matrix and recent test results |
| Operations | Who monitors, responds, changes and communicates? | Rota, runbooks, dashboards and incident record |
| Continuity | What must be restored, by when and to what point? | Successful restore exercise against targets |
Should infrastructure be internal, managed or co-managed?
Keep delivery internal when the infrastructure is strategically distinctive, the organization can recruit and retain the required skills, and demand justifies a dedicated team. Managed delivery can improve coverage where capabilities are standardized or needed around the clock. Co-management is often the practical choice: the business retains architecture, risk acceptance and product context while a provider handles defined operational tasks. The model should follow accountability and capability, not a blanket preference for outsourcing.
No model transfers executive accountability for business risk. Contracts should name the customer and provider responsibility for every recurring task and every incident decision. Include cloud-account ownership, privileged access, subcontractors, evidence access, data location, license ownership and exit assistance. A provider should not be the sole holder of domain registration, root identity, encryption keys or billing records. Those dependencies make an otherwise routine transition needlessly dangerous.
Which service levels actually matter?
Start with user journeys and tolerated impact, then choose service level indicators. The Google SRE guidance on service level objectives distinguishes observed indicators from objectives and contractual agreements. Availability alone is too blunt. A service may be reachable while requests fail, queues age or results become stale. Measure successful transactions, latency, correctness, freshness or durability as appropriate, and specify the observation window and exclusions.
Support terms should also cover acknowledgement, qualified response, restoration and communication. A five-minute acknowledgement does not promise a five-minute fix. Define severity using customer impact, not the component's name, and state who can declare a major incident. Establish maintenance windows, escalation paths and rules for clock pauses. Review service levels against incident evidence: permanently green dashboards alongside repeated user complaints usually indicate the wrong indicators or boundary.
How should security responsibility be divided?
Create a responsibility matrix for identity, device and workload credentials, network policy, hardening, encryption, logging, vulnerability handling, patching, backup, incident response and evidence retention. The NIST SP 800-53 control catalog is broad enough to help teams test coverage, but controls must be selected and tailored to context. Record inherited controls and customer-configured controls separately. A cloud platform's encryption capability does not prove that the customer enabled it correctly or protected the keys.
Require federated workforce identity, strong multi-factor authentication for privileged roles, short-lived administrative access where possible and auditable emergency access. Service identities should be unique, narrowly permitted and rotated without outage. Ask how the provider screens and removes staff access, how subcontractors are governed, where support sessions are logged and how security events reach the customer's incident process. Security reports are useful only if exceptions have owners and deadlines.
What should monitoring and incident response cover?
Monitoring should begin with user impact and then expose likely causes. The OpenTelemetry signals model describes traces, metrics and logs as complementary signals. Instrument request outcomes, latency distributions, saturation, dependency errors, deployment markers, capacity and security-relevant changes. Do not collect secrets or unnecessary personal data merely because telemetry storage is available. Specify retention, access and time synchronization so an investigation can reconstruct events.
Every alert needs a condition, owner, urgency, destination and response. Page for urgent customer impact or imminent objective loss; route slow trends into planned work. Incident terms should define command, technical work, customer communication, supplier escalation, evidence preservation and post-incident review. Test the contact path outside business hours. A runbook that has never been exercised and an escalation list with former employees are not operational controls.
| Operating signal | Why it matters | Review question |
|---|---|---|
| Objective attainment | Shows delivered user reliability | Are misses concentrated in one journey or dependency? |
| Change failure rate | Connects releases to instability | Which change classes need safer rollout? |
| Time to restore | Tests detection, authority and recovery | Where did responders wait or improvise? |
| Patch exposure | Shows age of relevant unresolved risk | Are exceptions risk-accepted and expiring? |
| Restore success | Proves usable recovery, not backup completion | Did restored service meet data and time targets? |
| Unallocated spend | Reveals assets without ownership | Should the resource be assigned or retired? |
What is the difference between backup and recovery?
Backup creates a recoverable copy; recovery restores an agreed business capability. Define recovery point objectives for acceptable data loss and recovery time objectives for acceptable interruption. The NIST contingency planning guide links business impact analysis, recovery strategies, plans, testing and maintenance. Include identity, encryption keys, configuration, infrastructure code and external integrations in exercises. Restoring a database is insufficient if applications cannot authenticate or reconcile later transactions.
Ask for recent restore evidence, not only successful backup-job percentages. Exercises should verify integrity, permissions, application behavior and business reconciliation in an isolated environment. Include ransomware scenarios, accidental deletion, bad deployment and provider-region failure where relevant. Record actual time, manual steps and missing dependencies. Recovery plans must change when architecture, ownership or suppliers change; an annual test of an obsolete topology produces false confidence.
How are infrastructure services priced?
Common models include fixed monthly scope, per-device or per-resource fees, consumption-based work and retained capacity with project rates. Compare total operating coverage, not headline unit price. Identify included onboarding, monitoring, patching, backup, incident hours, service reviews, tool licenses, cloud consumption, after-hours work and major changes. Demand assumptions for resource counts and support volume, plus a mechanism to approve growth. Unbounded pass-through cloud spend and vague project exclusions are frequent sources of surprise.
The contract should provide evidence rights, data portability, configuration and runbook ownership, vulnerability-notification duties, change approval, subcontractor disclosure, termination assistance and secure deletion. Define which artifacts are continuously maintained. Price the exit while negotiating the entry: credentials, inventories, infrastructure code, observability configuration and history should be transferable in usable formats. A low monthly fee can conceal expensive dependence if the customer cannot operate or move the service.
How should a provider be evaluated?
Give finalists the same representative scenario and ask them to show how they would discover dependencies, propose controls, respond to a failed release and prove recovery. Meet the people who will operate the account, not only sales staff. Review sample dashboards, runbooks, change records and post-incident reports with sensitive details removed. Validate certifications only for the legal entity, services and locations in scope. References should resemble your scale, technology and support model.
- Name the business service, owner, users and tolerated impact before requesting a quote.
- Require a responsibility matrix spanning routine work, change, security, incidents and recovery.
- Evaluate evidence quality: inventories, tests, telemetry, decision records and completed actions.
- Retain control of strategic accounts, identities, data, billing and transferable operating artifacts.
- Run a bounded transition with acceptance criteria rather than switching the entire estate at once.

Key takeaways
- Infrastructure services are an operating capability, not a hardware list.
- Service objectives should represent user outcomes and be paired with response and recovery terms.
- Customer and provider duties must be explicit for every control and operational decision.
- Recovery is proven by exercised restoration and reconciliation, not by backup status alone.
- Commercial comparison should include operating coverage, evidence access and exit readiness.
Frequently asked questions
Can one provider manage both cloud and on-premises infrastructure? Yes, if the responsibility model, tools and support skills cover both environments and their connecting dependencies. Does managed infrastructure include application support? Only when the scope says so; distinguish platform health from application ownership. How often should service reviews occur? Monthly is common for operating signals, with immediate reviews after major incidents and less frequent strategic reviews. What should happen first in a transition? Establish access, inventory, dependencies, current risk and rollback before transferring change authority.
Conclusion
The best answer to an infrastructure services FAQ is a service definition that can be operated and tested. Anchor the agreement in business journeys, named owners, measurable behavior, proportionate controls and recovery evidence. Whether delivery is internal, managed or shared, the organization should always know who can act, what success looks like and how it will regain control when conditions change.