These 26 AWS interview questions cover the work a senior AWS engineer does in production: IAM and multi-account structure, VPC networking, choosing between Lambda, ECS and EKS, DynamoDB data modeling, cost control and the Well-Architected Framework. They mirror how Ryz vets AWS engineers. Recruiters source people with real production AWS experience, structured NTRVSTA AI interviews probe how they reason through scenarios like these, and recruiters review each candidate before and after. The AI scores are advisory; people make the decisions.
Pick six to eight questions per interview and weight them by the role. Platform roles lean on networking and architecture; serverless API roles on Lambda and DynamoDB. Skip anything answerable by reciting a definition and push on the follow-up: "what broke the last time you did that?"
Good AWS answers start with constraints (traffic shape, compliance, team size, budget) and end with a trade-off. Pair the talk with the exercise below so you see them read a real IAM policy.
Requests are denied by default, and an explicit Deny in any applicable policy wins immediately. Organization guardrails (service control policies and resource control policies) cap what is possible in the account, permissions boundaries and session policies cap what a principal can do, and within those limits an Allow in an identity policy or a resource policy grants access. Cross-account access needs both sides: the caller's identity policy and the resource policy in the other account.
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::123456789012:role/orders-worker \
--action-names s3:GetObject \
--resource-arns arn:aws:s3:::acme-orders/exports/2026-09.csv
What a strong answer shows: they know the simulator only checks what you give it, so a bucket policy, KMS key policy or SCP can still deny a request the simulator allows.
Confirm which role is in use, then read the CloudTrail event. Causes: the policy grants the bucket ARN instead of bucket/*; the object is encrypted with SSE-KMS and the role lacks kms:Decrypt or the key policy excludes it; a bucket policy has an explicit deny (for example, requiring a specific VPC endpoint); a VPC endpoint policy restricts the bucket; or an SCP blocks the region.
What a strong answer shows: a systematic order of checks driven by CloudTrail, and awareness that KMS is a separate authorization decision.
Security groups attach to network interfaces, are stateful and allow-only, and can reference other groups. NACLs apply at the subnet, are stateless, evaluate numbered rules in order and support deny. Most designs rely on security groups referencing each other and leave NACLs near default.
What a strong answer shows: they mention ephemeral return ports when a NACL is tightened, which is the classic way people break traffic with NACLs.
Static keys leak and never expire. Humans should sign in through IAM Identity Center and get short-lived role sessions. CI systems should federate with OIDC: GitHub Actions requests a token, STS exchanges it for temporary credentials, and the trust policy pins exactly which repo and branch may assume the role.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:acme/payments-api:ref:refs/heads/main"
}
}
}]
}
What a strong answer shows: they scope the sub condition tightly and know that a wildcard like repo:acme/* lets any repo in the org deploy.
S3 with a lifecycle policy: Standard for the hot window, then transition to Glacier Instant Retrieval or Standard-IA, then Glacier Deep Archive for long retention, with an expiration matching the retention policy. Partition keys by date for Athena and compact small files, since colder classes bill minimum object sizes and durations.
What a strong answer shows: they consider retrieval patterns, minimum-duration charges and small-object overhead instead of picking the cheapest per-GB class.
Spreading across Availability Zones protects against the loss of a data center or a zone-level networking event. It does not protect against a regional service outage, a bad deploy, a corrupted database write replicated to the standby, or an IAM mistake.
What a strong answer shows: they separate availability from recoverability and mention backups or point-in-time restore for the failure modes replication copies faithfully.
SQS is a work queue: one consumer group pulls, retries and dead-letters messages. SNS fans one message out to many subscribers. EventBridge routes events by content rules across services and accounts. Kinesis Data Streams is an ordered, replayable log per shard with multiple independent readers.
What a strong answer shows: they pick based on ordering, replay and fan-out needs, and mention SNS-to-SQS fan-out as the common combined pattern.
Lambda fits spiky, event-driven traffic with short requests. ECS on Fargate fits steady services where you want containers without managing nodes. EKS makes sense when the organization already runs Kubernetes or needs its ecosystem.
What a strong answer shows: they ask about traffic shape, request duration, the team's existing platform and who will be on call before naming a service.
Enable partial batch responses on the event source mapping (ReportBatchItemFailures) and return only the failed message IDs, so successful messages are deleted. Add a redrive policy so poison messages reach a dead-letter queue, and keep the visibility timeout well above the function timeout.
import type { SQSEvent, SQSBatchResponse } from "aws-lambda";
export const handler = async (event: SQSEvent): Promise<SQSBatchResponse> => {
const batchItemFailures: SQSBatchResponse["batchItemFailures"] = [];
for (const record of event.Records) {
try {
await processOrder(JSON.parse(record.body));
} catch (err) {
console.error({ messageId: record.messageId, err });
batchItemFailures.push({ itemIdentifier: record.messageId });
}
}
return { batchItemFailures };
};
What a strong answer shows: they also make processOrder idempotent, because SQS standard queues deliver at least once.
Measure first: the Init Duration in the REPORT log line tells you whether cold starts are the problem. Then trim dependencies, initialize SDK clients outside the handler, and turn on SnapStart, which restores from a snapshot of the initialized environment. Provisioned concurrency covers a known baseline at a fixed cost.
What a strong answer shows: they know SnapStart has caveats around uniqueness and stale state captured in the snapshot, such as random seeds and open connections.
Start from access patterns. Store orders under the customer partition sorted by date, plus an item keyed by order ID. Status queries need care: a GSI keyed only on status puts every pending order in one partition.
PK SK notes
CUSTOMER#c_812 ORDER#2026-09-14#o_5521 query by customer, date range on SK
ORDER#o_5521 META direct get by order id
GSI1 (sparse, only set while status is PENDING)
GSI1PK = PENDING#3 (shard 0-9 picked from hash of order id)
GSI1SK = 2026-09-14T10:02Z#o_5521
What a strong answer shows: they use a sparse index and write sharding for the low-cardinality status pattern, and they ask how many pending orders exist before over-engineering it.
On-demand still has per-partition limits and scales from previously observed peaks, so a sudden jump or a hot key throttles. Find hot keys with CloudWatch Contributor Insights, spread writes with better keys, retry with jitter, and pre-warm throughput before planned events.
What a strong answer shows: they diagnose a hot partition versus table-level capacity instead of assuming on-demand means unlimited.
Every concurrent execution opens its own connection. Put RDS Proxy in front to pool and multiplex, cap reserved concurrency on the function and open connections outside the handler so warm invocations reuse them.
What a strong answer shows: they know proxies can "pin" sessions (for example with session-level settings or prepared statements), which removes the multiplexing benefit.
Public subnets per AZ hold the ALB and NAT gateways; private application subnets host ECS tasks or EC2; isolated data subnets hold Aurora and ElastiCache with no route to the internet. Use one NAT gateway per AZ. Size the CIDR generously and avoid overlap with other VPCs and on-premises ranges.
What a strong answer shows: they plan IP space for growth and add VPC endpoints for S3, DynamoDB and ECR from day one.
Enable VPC Flow Logs and query them in Athena to find top talkers through the NAT. Usual suspects: ECR image pulls, S3 traffic and AWS API calls going through the NAT instead of endpoints. Add free gateway endpoints for S3 and DynamoDB and interface endpoints for high-volume services like ECR, and check for cross-AZ routing to another AZ's NAT.
What a strong answer shows: they diagnose from flow data before adding endpoints, since interface endpoints have their own hourly and per-GB costs.
Peering is simple for a few VPCs but non-transitive and needs non-overlapping CIDRs. Transit Gateway is the hub for many VPCs, VPNs and Direct Connect with central routing. PrivateLink exposes one service behind a Network Load Balancer to consumers without routing whole networks together, and works even when CIDRs overlap.
What a strong answer shows: they default to exposing services, not networks, when the need is "call this API".
KMS issues a data key; the service encrypts data locally with it and stores the encrypted copy of the key alongside the data. Cross-account use requires the key policy to allow the other account and that account's IAM policy to grant kms:Decrypt; AWS managed keys can't be shared this way, so cross-account workloads need customer managed keys.
What a strong answer shows: they know key policies are the primary control and that a key policy without the account root principal can lock everyone out.
IMDSv2 requires a session token obtained with a PUT request, which blocks most SSRF attempts to steal instance credentials. Old SDKs and agents that only speak IMDSv1 break, and containers on EC2 need a hop limit of 2 to reach the metadata service.
aws ec2 modify-instance-metadata-options \
--instance-id i-0abc1234def567890 \
--http-tokens required \
--http-put-response-hop-limit 2
What a strong answer shows: they roll it out by first checking the MetadataNoToken CloudWatch metric to find IMDSv1 callers.
Deactivate the key immediately, then search CloudTrail for every call made with it: new IAM users, roles, keys, EC2 instances in unused regions, Lambda functions or changes to bucket policies. Remove any persistence and check GuardDuty findings. Then replace the key with role-based credentials.
What a strong answer shows: they hunt for persistence across all regions, not just delete the key and move on.
Use AWS Organizations with a Control Tower landing zone: accounts for security, log archive, networking and each workload per environment. Apply SCPs for guardrails (deny disabling CloudTrail, restrict regions) and centralize identity in IAM Identity Center.
What a strong answer shows: they treat accounts as the blast-radius and billing boundary and plan account vending so teams aren't blocked waiting on tickets.
Start with RTO and RPO per service, because they decide the pattern: backup and restore, pilot light, warm standby or active-active. Then address data, the hard part: Aurora Global Database, DynamoDB global tables and S3 replication fail over differently. Keep the failover path free of dependencies on the failed region, and rehearse it.
What a strong answer shows: they push back on blanket active-active, discuss conflict resolution for multi-writer data, and insist on regular failover drills.
Get visibility first: cost allocation tags, the Cost and Usage Report and a breakdown by service and team. Then delete idle resources, rightsize with Compute Optimizer, move suitable workloads to Graviton, fix data transfer, and only then buy Compute Savings Plans sized to the steady baseline.
What a strong answer shows: they buy commitments last, after rightsizing, and they set up ongoing ownership so costs don't drift back.
Scope it to one workload with the people who run it, and walk the six pillars: operational excellence, security, reliability, performance efficiency, cost optimization and sustainability. The output is a short list of high-risk issues with owners and dates, not a slide deck.
What a strong answer shows: they can name a real finding from a past review and what it took to fix it.
Check every quota in the path: account-level Lambda concurrency, API Gateway throttling, DynamoDB capacity and any third-party APIs downstream. Load test a production-like environment, request quota increases early and cache read-heavy endpoints. Decide in advance what to shed.
What a strong answer shows: they find the weakest dependency (often a database or external API) instead of only scaling the parts AWS scales for you.
Assume at-least-once delivery everywhere. Make consumers idempotent with an idempotency key stored in DynamoDB using a conditional write, or use Powertools for AWS Lambda's idempotency utility. Where order matters, use SQS FIFO message groups or Kinesis partition keys, plus versions so consumers discard stale updates.
What a strong answer shows: they call "exactly once" a property of the consumer's design, not something the queue provides end to end.
Containerize the monolith as-is first and run it on ECS behind the same ALB, so deploys become repeatable before anything else changes. Then extract services with the strangler pattern via ALB listener rules, keeping rollback a routing change away.
What a strong answer shows: they sequence the work to reduce risk and avoid splitting a shared database across services in one step.
"Action": "*" or AdministratorAccess for application roles "to get it working" and never comes back.Give the candidate a repository (two to three hours, take-home or pairing) with Terraform or CDK for a webhook ingestion service: API Gateway, a Lambda writing to SQS, a worker Lambda and a DynamoDB table. Seed it with problems: wildcard IAM, no dead-letter queue, a visibility timeout shorter than the worker timeout, a scan-based query. Ask them to fix the three most serious issues and write a short note on next steps and the cost at 10 times today's volume.
Ryz can introduce senior AWS engineers who have already been through this kind of vetting. They are the top 1% of the candidates we interview, they join your team, accounts and repos, and they work within ±1h of US time zones so reviews and incident calls happen the same day. See how the vetting process works, or start from our AWS engineer job description.
They show familiarity with the service catalog, which helps, but they don't show judgment under real constraints.
Use code review instead of provisioning. Infrastructure code, IAM policies and CloudTrail events need no credentials to discuss. For hands-on work, a sandbox account with a strict budget and SCPs is enough.
Mid-level engineers build well with individual services. Senior engineers think in accounts, blast radius, failure modes and cost, and they can say no to an architecture that is more complex than the team can operate.
Questions we didn't answer? Email info@ryzlabs.com.