Key Insights From AWS

Explore top LinkedIn content from expert professionals.

Summary

Key insights from AWS reveal important lessons about cloud reliability, resilience, partnership strategies, AI-driven automation, and cost management. AWS—Amazon Web Services—is a leading cloud computing platform that powers many of the world’s online services, and its operational reports, ecosystem developments, and best practices offer valuable guidance for both technical and business leaders.

  • Build for resilience: Design your systems to withstand failures by considering multi-region setups and understanding how interdependent services can impact each other.
  • Embrace automation carefully: Use automated tools for detection and recovery, but always ensure human oversight to catch issues that technology may miss or inadvertently create.
  • Monitor and manage costs: Treat cloud cost management as an ongoing process, regularly reviewing workloads, billing structures, and operational practices to prevent unnecessary expenses.
Summarized by AI based on LinkedIn member posts
  • View profile for Wias Issa

    CEO @ Ubiq | Cybersec Exec | Built Businesses From Zero to Scale | Global P&L, Product & Enterprise GTM

    6,942 followers

    The detailed incident report from AWS is now public, and it’s well worth a read (link in comments). Here’s a distilled summary of what went wrong, and what tech leaders should take away. What happened: 1️⃣ A race condition in the DNS management system serving DynamoDB in US-EAST-1 led to endpoint resolution failures. 2️⃣ That dominant database service failure cascaded: new EC2 launches failed due to lease-management issues (on which EC2 depends) and network components suffered health-check failures that rippled across load balancers. 3️⃣ The impact was global. Apps and critical services relying on AWS saw outages, degraded performance, or intermittent failures. Why this matters: 1️⃣ Concentration risk: Even for a hyperscale provider like AWS, a failure in one region and one service (DynamoDB DNS) can cascade globally, turning a “cloud issue” into a business continuity event. 2️⃣ Complex interdependencies: The issue wasn’t just database DNS; it propagated into compute, networking, automation, and customer-facing systems. We often design for failure at one layer but underestimate coupling across layers. 3️⃣ Recovery complexity = resilience risk: Recovery isn’t just restarting services; it’s clearing backlogs, restoring state, and ensuring downstream systems don’t remain impaired. My perspective/takeaways: 1️⃣ Design for worst-case provider failure. Not just “an AZ down,” but “core service in region down” and the ripple effects. 2️⃣ Visibility and dependency mapping matter, so know what services your stack depends on, and how managed service failures might cascade. 3️⃣ Recovery orchestration is as vital as fault tolerance, so plan for backlog recovery, state cleanup, and cross-team communication. 4️⃣ Cloud-vendor resilience is not infinite, and shared failure domains persist even in hyperscale clouds. Plan for multi-region or cross-provider fallback and clear internal recovery roles. 5️⃣ Executive mindset and risk alignment. For C-suites, this is a reminder: infrastructure risk is business risk. Discuss cloud-failure modes at the board table, not just application risk. What this isn't about: This isn’t about blaming AWS. The lesson is that even the largest provider can experience a systemic failure, and we can all learn from these experiences. And... it's always DNS 😉

  • View profile for Hagay Lupesko

    Senior Vice President, AI Inference @ Cerebras Systems

    17,687 followers

    🚨 Lessons from the AWS us-east-1 outage on Oct 19 🚨 A single low-level DNS automation bug in DynamoDB propagated into a massive multi-service, region-wide failure lasting over 14 hours. So much is built on AWS that for a while it seemed as if the entire internet was down... Some interesting details from AWS’s postmortem 👇 🧩 Root cause: A race condition in DynamoDB’s DNS automation deleted its own regional endpoint. Damn! ⚡ Mitigation: AWS engineers identified and fixed the root cause in just over 2 hours. That's impressive, given AWS's scale. Kudos to the AWS on-calls! 🏗️ AWS runs on AWS: The DynamoDB failure cascaded to EC2, NLB, Lambda, Redshift, ECS, EKS, SQS, and many other services. It’s amazing to see how deep the rabbit hole goes and how much AWS is built on top of AWS! 🤖 Automation paradox: The very automation meant to speed recovery caused a “congestive collapse” in EC2’s recovery workflow. It was only resolved once human on-calls intervened to manually throttle and clear queues. There’s still hope for humanity! 🙌 💭 The bigger lesson: Service outages, in particular at hyper scale, are inevitable. If something can fail, it will fail - good old Murphy's law! So how do you protect against the next cloud outage? ✅ Redundancy: Architect your service for multi-region resiliency. Active-active or active-passive failover, pick your poison. If your service is not architected this way, you're vulnerable! ⚙️ Detection & mitigation: Real-time metrics, fast alerts, and a world-class on-call culture are all key. Having all of that is what enabled AWS to detect and fix the root cause within hours. 📚 Learning from failures: AWS’s postmortem is a masterclass in rigorous incident analysis, and it was powered by Amazon's notorious Correction of Errors (CoE) process. Every engineering org should have such process in place. This is how you ensure continuous improvements!

  • View profile for Neeti Gupta

    PhD Candidate at University of Cambridge. Founder of AI Partnerships. Former Microsoft, Meta, Amazon, GE Healthcare, VMware, Broadcom | New Business Development

    17,089 followers

    AWS Ecosystem Analysis: Partnership Strategy and Market Impact Based on analysis of interviews with AWS partnership leaders, here are key findings about AWS's ecosystem approach: Partnership-Centric Model - AWS operates with dual focus on customers and partners, viewing partners as force multipliers providing industry expertise. Internal KPIs align with customer growth metrics, reflecting that partner success correlates with AWS performance. AWS Marketplace Performance - The Marketplace has evolved from AMI distribution to a comprehensive platform serving 300,000+ active customers, with all top 1,000 AWS customers utilizing it. Key metrics (let me know if these numbers have changed): - 27% higher win rates for marketplace transactions - 80% increase in deal values - 40% faster sales cycles - Procurement reduced from months to days AWS implements "Marketplace Everywhere," integrating purchasing into service consoles (EC2, RDS, EKS) and providing API access for custom storefronts. The goal is positioning AWS Marketplace as the primary enterprise IT procurement channel. Partnership Framework - AWS works with technology partners (ISVs), consulting partners (GSIs, Advisory, Born-in-the-Cloud), and channel partners (Distributors, Resellers). Partner progression model: - Validate: Technical solution verification - Spotlight: Enhanced resources and management attention - Endorsed: Joint go-to-market activities Partners are directed toward targeted approaches focusing on specific customer profiles and industry segments. AWS invests in global marketplace localization for currency, language, and regulatory compliance. AI Integration - AWS characterizes AI as a foundational industry shift, integrating capabilities across all solution areas. Partners extend core AI services like Bedrock with domain-specific applications. The company encourages building on their AI stack immediately, calling out unprecedented market demand. Co-selling Operations - Partnership leaders must demonstrate quantifiable value through metrics including deal velocity and success rates. Sales teams are incentivized to co-sell with partners, and partners are encouraged to consistently register opportunities (e.g., through ACE) to gain visibility and establish patterns of success with AWS field teams. Performance Measurement - AWS evaluates partners across four dimensions: market presence expansion, partner-generated deal flow, product capability enhancement, and customer retention improvement. Centralized tools enable real-time ROI tracking and adoption monitoring. Market Implications - AWS's ecosystem strategy demonstrates how cloud platforms scale through strategic partnerships. The marketplace model represents a shift toward platform-mediated procurement, establishing new standards for enterprise technology acquisition. Thoughts or something I missed or got wrong? Feel free to comment below. #AWS #CloudComputing #Partnerships #EnterpriseStrategy

  • View profile for Matt Wood
    Matt Wood Matt Wood is an Influencer

    Chief AI & Technology Officer, AWS

    88,918 followers

    AI field note: AI is moving from faster answers to work that carries forward. An answer can make a single task faster, but when the work builds on what came before, each step improves the next. Our new launches carry work forward between steps, so one stage can become the start of the next; between sessions, so what the system learns today is there tomorrow. 🔒 Continuity in software and security. Software is never really finished, and keeping pace has always required constant effort. AI changes that by giving improvement a tailwind, so codebases can keep moving forward instead of relying on episodic cleanup. AWS Continuum works the lifecycle of threats and vulnerabilities. It prioritizes weaknesses, validates which ones are genuinely exploitable and verifies patches all within guardrails you set. It can also build threat models straight from your design docs or source code. Release Management in AWS DevOps Agent reviews whether a change is ready and tests it before it ships, so teams can catch a breaking change before it reaches production. AWS Transform now runs continuously behind your coding agents, finding tech debt, fixing it, validating the fix, and keeping packages up to date so codebases stay current as the work moves forward. Together, they create a loop: find what needs attention, improve it, validate it, ship it, and feed what was learned back in, so the next cycle starts from what the last one taught. ⚡️ Continuity for agents. As more work is carried by agents, they need continuity too, not just for one run but across the whole lifecycle. AgentCore Harness, Web Search, Policy Guardrails, and Optimization help agents move from idea to production to improvement: assemble the agent, connect to useful tools, govern what it can do, observe where it succeeds or drifts, and improve over time. 📚 Continuity in understanding. AWS Context automatically builds a knowledge graph from your data. It works out how your tables, documents, business rules, and domain knowledge relate, and makes that available to every agent, with built-in governance so each one only sees what it is allowed to. Context also learns as agents use it. When one agent finds the right way to answer a question, every other agent can use that same path. Bedrock Managed Knowledge Base handles unstructured retrieval automatically, pulling from sources like S3, SharePoint, Confluence, and Google Drive, and plugs into Context so agents can search a unified view of your data. 🚀 Continuity in work. Amazon Quick's new autonomous agents run in the background with their own expertise and access. They can process orders as they come in, or watch a CRM, an inbox, and Slack to draft follow-ups, flag risks, and suggest the next step, all with no code required. Quick also uses the same agentic search as AWS Context. When work starts fresh each time, you get efficiency. When it carries forward, decisions and history intact, that efficiency compounds into reinvention.

  • View profile for Shishir Khandelwal
    Shishir Khandelwal Shishir Khandelwal is an Influencer

    Staff Engineer at PhysicsWallah

    21,157 followers

    Alongside building resilient, highly available systems and strengthening security posture, I’ve been exploring a new focus area, optimising cloud costs. Over the last few months, this has led to some clear lessons for me that are worth sharing. 1. Compute planning is the foundation. Standardising on machine families and analysing workload patterns allows you to commit to savings plans or reserved instances. This is often the highest ROI move, delivering big savings without actually making a lot of technical changes. 2. Account structures impact cost. Multiple AWS accounts improve governance and security but make it harder to benefit from bulk discounts. Using consolidated billing and commitment sharing across accounts brings the efficiency back. 3. Kubernetes compute checks are important. Nodes in K8s are often over-provisioned or underutilised. Automated rebalancing tools help, as does smart use of spot instances selected for reliability. On top of this, workload resizing during off hours, reducing CPU and memory when demand is low, delivers direct and recurring savings. 4. Watch for operational leaks. Debug logs on CDNs and load balancers, once useful, often stay enabled long after issues are fixed. They quietly pile up costs until someone takes notice. 5. Right-sizing is a continuous process. Urgent projects often lead to overprovisioned instances for anticipated load that never fully arrives. Monitoring and regular reviews are the only way to keep infrastructure aligned with reality. The real win in cloud cost optimisation comes from treating it as a continuous practice, not a one-off project. Small inefficiencies compound fast, so important to be on the lookout! #CloudCostOptimization #AWS #Kubernetes #DevOps #CloudInfrastructure #RightSizing #WorkloadManagement #SavingsPlans #SpotInstances #CloudEfficiency #TechInsights #CloudOps #CostManagement #CloudBestPractices

  • View profile for Vishakha Sadhwani

    Sr. Solutions Architect at Nvidia | Ex-Google, AWS | EB1-A Recipient || Opinions, my own ||

    177,158 followers

    The AWS downtime this week shook more systems than expected - here’s what you can learn from this real-world case study. 1. Redundancy isn’t optional Even the most reliable platforms can face downtime. Distributing workloads across multiple AZs isn’t enough.. design for multi-region failover. 2. Visibility can’t be one-sided When any cloud provider goes dark, so do its dashboards. Use independent monitoring and alerting to stay informed when your provider can’t. 3. Recovery plans must be tested A document isn’t a disaster recovery strategy. Inject a little chaos ~ run failover drills and chaos tests before the real outage does it for you. 4. Dependencies amplify impact One failing service can ripple across everything. You must map critical dependencies and eliminate single points of failure early. These moments are a powerful reminder that reliability and disaster recovery aren’t checkboxes .. They’re habits built into every design decision.

  • View profile for Danny Steenman

    Helping startups build faster on AWS while controlling costs, security, and compliance | Founder @ Towards the Cloud | Freelancer

    11,478 followers

    I recently completed a client's AWS infrastructure audit. The issues that uncovered are surprisingly common. Here's what I found: 𝟭. 𝗨𝗻𝗲𝗻𝗰𝗿𝘆𝗽𝘁𝗲𝗱 𝗘𝗕𝗦 𝗩𝗼𝗹𝘂𝗺𝗲𝘀   Data at rest was not encrypted, posing a significant security risk. 𝟮. 𝗖𝗹𝗼𝘂𝗱𝗧𝗿𝗮𝗶𝗹 𝗗𝗶𝘀𝗮𝗯𝗹𝗲𝗱   The account lacked crucial audit logs, limiting visibility into account activities. 𝟯. 𝗣𝘂𝗯𝗹𝗶𝗰 𝗦𝟯 𝗕𝘂𝗰𝗸𝗲𝘁𝘀   Several S3 buckets were publicly accessible, potentially exposing sensitive data. 𝟰. 𝗦𝗦𝗛 (𝗣𝗼𝗿𝘁 𝟮𝟮) 𝗢𝗽𝗲𝗻 𝘁𝗼 𝘁𝗵𝗲 𝗪𝗼𝗿𝗹𝗱   Unrestricted SSH access increased the attack surface unnecessarily. 𝟱. 𝗩𝗣𝗖 𝗙𝗹𝗼𝘄 𝗟𝗼𝗴𝘀 𝗗𝗶𝘀𝗮𝗯𝗹𝗲𝗱   Network traffic insights were missing, hampering security analysis capabilities. 𝟲. 𝗗𝗲𝗳𝗮𝘂𝗹𝘁 𝗩𝗣𝗖 𝗦𝘁𝗶𝗹𝗹 𝗶𝗻 𝗨𝘀𝗲   The default VPC was being used, often lacking proper segmentation and security controls. These findings aren't unusual. Many organizations, from startups to enterprises, overlook these aspects of AWS security and best practices. That's why doing regular AWS account audits are crucial. They help identify potential vulnerabilities before they become problems. 𝗞𝗲𝘆 𝘁𝗮𝗸𝗲𝗮𝘄𝗮𝘆𝘀 𝗮𝗻𝗱 𝘀𝗼𝗹𝘂𝘁𝗶𝗼𝗻𝘀: 1. Encrypt data at rest: Enable default EBS encryption at the account level. 2. Implement comprehensive logging: Enable CloudTrail across all regions and set up alerts. 3. Restrict public access: Use S3 Block Public Access at the account level and audit existing buckets. 4. Use modern, secure access methods: Implement AWS Systems Manager Session Manager instead of open SSH. 5. Enable network monitoring: Turn on VPC Flow Logs and set up automated analysis. 6. Design your network architecture intentionally: Create custom VPCs with proper security controls. By addressing these common issues, you significantly enhance your AWS security posture. It's not about perfection, but continuous improvement. When's the last time you audited your AWS environment?

  • View profile for Sarthak Rastogi

    AI engineer | Posts on agents + advanced RAG | Experienced in LLM research, ML engineering, Software Engineering

    30,762 followers

    Amazon Web Services (AWS) just shared their real-world lessons from building 1000s of agents at Amazon. Most teams still evaluate like it’s a single LLM call. But agents aren’t just outputs anymore. They’re systems. - multi-step reasoning - tool selection + execution - memory retrieval - multi-agent coordination If you only evaluate the final response, you miss why things break. AWS just shared how they evaluate agents internally, and it’s a much more practical approach. They break evaluation into 3 layers: 1. Bottom layer: model performance (latency, accuracy, cost) 2. Middle layer: components (intent detection, planning, tool use, memory) 3. Top layer: final outcome (task success, UX, safety) Instead of asking “did it answer correctly?”, you can localize failure: - planning score: was the task decomposition valid? - tool selection accuracy: was the right capability invoked? - tool call error rate: did execution fail or inputs break? - grounding / faithfulness: did reasoning stay consistent with context? - multi-turn coherence: did state drift over time? Another key shift: evaluating trajectories, not just outputs. Agent traces (reasoning steps, tool calls, intermediate states) become first-class evaluation artifacts. - replay traces against new versions -- regression testing - generate synthetic eval datasets from production logs - benchmark tool-use sequences, not just answers This is especially critical in multi-agent setups, where failure modes include: - poor task decomposition by the orchestrator - incorrect agent assignment - communication breakdown between agents - inconsistent aggregation of results Finally, one thing that consistently shows up in production: You need human-in-the-loop, not as a fallback, but as part of the eval loop. - calibrate LLM-as-a-judge - audit edge-case trajectories - validate reasoning quality, not just correctness ♻️ Share it with anyone who is building AI agents in production :) I share tutorials on how to build + improve AI apps and agents, on my newsletter 𝑨𝑰 𝑨𝒈𝒆𝒏𝒕 𝑬𝒏𝒈𝒊𝒏𝒆𝒆𝒓𝒊𝒏𝒈: https://lnkd.in/gaJTcZBR Link to article: https://lnkd.in/eFT45Gsq #AI #AIAgents #LLMs

  • View profile for Charisma DeLeon Island CISSP

    Data & AI Governance, Risk & Compliance | Multi-Cloud Security Architect | Cybersecurity Advisor | Public Speaker | Designing Secure & Compliant Enterprise Solutions

    5,883 followers

    As a former AWS Technical Delivery Manager, I taught hundreds of customers how to migrate their workloads to AWS. Last week, I spent a few days working with individuals on a migration project, and I'm sharing a few tips below. First, 𝐀𝐖𝐒 𝐀𝐩𝐩𝐥𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝐲 𝐒𝐞𝐫𝐯𝐢𝐜𝐞 (𝐀𝐃𝐒) removes the guesswork with EC2 recommendations to run your workloads to plan migrations with AWS Migration Hub by:  • Gathering Server and DB inventory for Database Migration Service.  • Server utilization data to generate rightsized EC2 instances.  • Map network communication patterns to understand application dependencies and group servers together.  • Export processes are running on the servers with agents installed. Second, 𝐀𝐖𝐒 𝐃𝐚𝐭𝐚𝐛𝐚𝐬𝐞 𝐌𝐢𝐠𝐫𝐚𝐭𝐢𝐨𝐧 𝐒𝐞𝐫𝐯𝐢𝐜𝐞 (𝐃𝐌𝐒) makes it easy to securely assess, convert, and automate the migration of your databases and analytics workloads with network controls and real-time visibility. DMS minimizes operational disruptions to your applications by keeping source systems fully operational until the migration is complete. Third, 𝐀𝐖𝐒 𝐌𝐢𝐠𝐫𝐚𝐭𝐢𝐨𝐧 𝐇𝐮𝐛 is a centralized platform that enables you to monitor your migration from planning to end-to-end execution, providing automated recommendations to accelerate your transformation. What I really like is these services are included in the Free and Paid plan tiers, allowing SMBs with AWS credits to evaluate their workloads for migration and modernization. 𝑾𝒆 𝒔𝒑𝒆𝒏𝒕 𝒍𝒆𝒔𝒔 𝒕𝒉𝒂𝒏 $10  to gather server information, EC2 recommendations, and test cutover. For 𝐀𝐈 𝐰𝐨𝐫𝐤𝐥𝐨𝐚𝐝𝐬 𝐚𝐧𝐝 𝐭𝐡𝐞 𝐆𝐏𝐔-𝐚𝐬-𝐚-𝐬𝐞𝐫𝐯𝐢𝐜𝐞 𝐦𝐚𝐫𝐤𝐞𝐭, analysts project that small and medium-sized businesses will allocate more than half of their technology budgets to cloud services. With the cloud migration market expected to grow from $232B to $806B by 2029 (+28%), SMBs are leading the charge, especially those investing in AI, AIOps, and DevOps to modernize faster. Starting in November, 𝐀𝐖𝐒 𝐓𝐫𝐚𝐧𝐬𝐟𝐨𝐫𝐦 takes things a step further as the first agentic AI service developed to accelerate enterprise modernization by deploying specialized AI agents to automate complex tasks, such as assessments, code analysis, refactoring, decomposition, dependency mapping, validation, and transformation planning, thereby dramatically reducing project timelines. The service helps reduce both modernization costs and ongoing maintenance expenses while identifying opportunities to eliminate legacy licensing costs for large enterprises. AWS Transform is the next leap bringing agentic AI into migration and modernization. If you’ve tested any of these new AI-driven migration tools, I’d love to hear your experience.

  • View profile for Amrit Jassal

    CTO at Egnyte Inc

    2,914 followers

    At the recently concluded AWS re:Invent, Werner Vogels shared some critical lessons that are universal to improving architecture and processes within Engineering teams across the board. As systems inevitably grow in complexity over time, he suggests embracing evolution and building with simplicity and manageability in mind from day one. Some of the key lessons about managing complexities that were worth noting include: 1. Make evolvability a requirement: Design systems knowing they will change. Prioritize flexibility and anticipate future needs. For instance, Amazon S3 has a simple API that has remained consistent while the underlying architecture has undergone radical transformations to accommodate growth and new features. 2. Break complexity into pieces: Decompose systems into smaller, manageable components with well-defined interfaces. This allows for independent scaling, evolution, and maintenance. Amazon CloudWatch has evolved from a simple service to a collection of microservices to improve functionality and address engineering challenges. 3. Align your organizations to your architecture: Structure teams to mirror the architecture of your systems. This promotes ownership, clear responsibilities, and efficient development. It is important for teams to own their work and for leaders to foster a sense of agency and urgency. 4. Organize into cells: Divide systems into isolated cells to limit the impact of failures and disturbances. This approach enhances reliability and simplifies operational management. Vogels explains how various AWS services like CloudFront and Route 53 utilize cell-based architectures. 5. Design predictable systems: Minimize uncertainty by designing systems with predictable behavior. Ensure consistent processing and avoid spikes or bottlenecks. 6. Automate complexity: Automate everything that doesn't require human judgment. This frees up resources and reduces the risk of human error. AWS, for instance, leverages automation extensively, particularly in security, with automated threat intelligence and agent-based workflows for support tickets. A link to the complete session is available here: https://lnkd.in/gxWquATs

Explore categories