Reviewing Progress Regularly

Explore top LinkedIn content from expert professionals.

  • View profile for Marcus Chan

    I help B2B founders & owners build a sales team that runs without them | Deals move in 30 days, then a repeatable system that keeps them closing | $195M ex-Fortune 500 exec | WSJ + USA Today bestseller | 700+ clients

    102,466 followers

    I just watched an AE lose a $1.2M deal after running a "successful" product trial that the prospect LOVED. After 8 weeks of work, the CFO killed it with five words: "Let's try our current vendor." After analyzing 200+ enterprise sales cycles at companies including Salesforce, HubSpot, Thomson Reuters, and Workday, I've identified the exact framework that separates 80%+ trial conversion rates from the industry average of 30%. The psychological shift required… Stop treating trials as product demos and start treating them as RISK ELIMINATION EXERCISES. After being promoted 12 times and hitting #1 in every role before leading a 110-person team to $190M+ annually, I've developed a framework that's transformed how top companies run trials. THE 5 POINT TRIAL QUALIFICATION SYSTEM: 1. 𝗣𝗥𝗢𝗕𝗟𝗘𝗠 𝗩𝗔𝗟𝗜𝗗𝗔𝗧𝗜𝗢𝗡 Ask these 3 questions before any trial: → "What happens if you don't solve this in 90 days?" (quantify impact)  → "How have you tried solving this before?" (establishes solution gap)  → "Who else is affected?" (identifies stakeholders) These eliminate 68% of unqualified trials before they start. 2. 𝗦𝗨𝗖𝗖𝗘𝗦𝗦 𝗗𝗘𝗙𝗜𝗡𝗜𝗧𝗜𝗢𝗡 Document these 4 criteria: → Technical requirements (features that must work)  → Business metrics (quantifiable outcomes)  → Timeline requirements (implementation speed)  → User adoption requirements (usage patterns) Get confirmation: "If we demonstrate [criteria], you'd move forward with purchase by [date]. Correct?" 3. 𝗦𝗧𝗔𝗞𝗘𝗛𝗢𝗟𝗗𝗘𝗥 𝗠𝗔𝗣𝗣𝗜𝗡𝗚 Create a "Decision Matrix" for: → Technical buyers (every trial user)  → Economic buyers (CFO/budget holder)  → Political influencers (who can kill it)  → Current solution advocates (status quo beneficiaries) Document each person's personal win/loss if change happens. 4. 𝗣𝗥𝗘-𝗧𝗥𝗜𝗔𝗟 𝗔𝗚𝗥𝗘𝗘𝗠𝗘𝗡𝗧 Have legal review BEFORE starting: "We typically have legal review the agreement structure ahead of time so there are no surprises and to save us both time so we can hit the deadline of December 1st you set. Would you be open to this during the trial?" 5. 𝗖𝗨𝗥𝗥𝗘𝗡𝗧 𝗩𝗘𝗡𝗗𝗢𝗥 𝗦𝗧𝗥𝗔𝗧𝗘𝗚𝗬 Ask: → "Have you discussed these challenges with your current vendor?"  → "What was their response?"  → "What specific capabilities do they lack?" Document these to prevent the "let's try our current vendor" objection. RESULTS from this framework: ✅ Trial conversion: 32% to 83% in 60 days  ✅ Average deal size: +40%  ✅ Sales cycle: -37%  ✅ Forecast accuracy: +92%  ✅ Time on unsuccessful trials: -43% — Hey Sales Leaders! Want to see how we can install these kinds of results into your org? Go here: https://lnkd.in/ghh8VCaf

  • View profile for Caleb Mellas

    Engineering leader at Olo · Previously Wisely (acquired by Olo, $187M) · Writing to 33k+ engineers at Level Up Software Engineering 🚀 · Pioneering AI-agent-first SDLC + product building

    37,595 followers

    “Oof! What did I even do this week?!” ...Meetings?!👇🏼 As engineers we tend to focus on how much code we wrote, and how many Jira tickets we moved to done. But as we move into more senior+, tech lead and management positions, we start writing less and less code. Less coding / tickets leaves us feeling like we didn’t get anything done. 🤪 But are tickets/coding the only value we contribute to our team? I’ve wrestled a lot with the feeling of getting nothing done as I transitioned into a senior+ engineer, and tech lead... Slowly but surely it’s become clear to me. The more senior we become, the more important it is that we multiply our efforts and help others level up. 🚀 If I improve my coding by 10%, the team gets minimally better. 🚀 If I lean into being a force-multiplier and help 3-5 others level up, the whole team just got massively more productive and effective. 🚀🚀🚀🚀🚀 My efforts are compounded. Suddenly all these force-multiplying activities make sense! If I: – Give through, timely and helpful code reviews – Mentor junior engineers in our patterns, and systems – Tackle bugs or pain points that are plaguing the team – Write and review technical specs to ensure good design – Help estimate effort in roadmapping and planning meetings I’m now influencing and moving things forward on a whole different level. My productivity is actually improved – if we measure it by value contributed to the team and the business goals. 🙌🏼 But back to feeling unproductive... A lot of wins are harder to measure and leave us feeling unproductive. One way to help your brain is keeping a daily + weekly summary journal doc. 🧠 Daily: Write 3-5 bullet points of ways you helped the team move forward. Weekly: Roll those up into a short summary of wins/learnings. Monthly: Add bigger wins to your brag doc You can even send that weekly doc to your manager and product partners. It’s a great way to manage up and give visibility into your work. It also helps them point out areas you might think are valuable, but aren’t as much, or areas you should lean into even more. Personally, I’m on a journey to re-think that I’m still contributing huge value even without coding as much. Days where I don’t code are still successful if I’ve been a force-multiplier. 💪🤝 - - - - - - - - - - - - - - - - - - - - - - - - - - - - How about you? How do you think about contributing as a senior IC / manager when you aren’t coding as much? 🙋♀️🙋♂️ - - - - - - - - - - - - - - - - - - - - - - - - - - - - If you liked this post, you’ll probably love my weekly newsletter: https://lnkd.in/e8d5ymr3 👉 Follow along as I share everything I’ve learned about becoming a #fullstackengineer and leveling up into a #senior+ engineer and #techlead at a hyper-growth #startup.

  • View profile for Sohrab Rahimi

    Director, AI/ML Lead @ Google

    24,316 followers

    Evaluating LLMs is hard. Evaluating agents is even harder. This is one of the most common challenges I see when teams move from using LLMs in isolation to deploying agents that act over time, use tools, interact with APIs, and coordinate across roles. These systems make a series of decisions, not just a single prediction. As a result, success or failure depends on more than whether the final answer is correct. Despite this, many teams still rely on basic task success metrics or manual reviews. Some build internal evaluation dashboards, but most of these efforts are narrowly scoped and miss the bigger picture. Observability tools exist, but they are not enough on their own. Google’s ADK telemetry provides traces of tool use and reasoning chains. LangSmith gives structured logging for LangChain-based workflows. Frameworks like CrewAI, AutoGen, and OpenAgents expose role-specific actions and memory updates. These are helpful for debugging, but they do not tell you how well the agent performed across dimensions like coordination, learning, or adaptability. Two recent research directions offer much-needed structure. One proposes breaking down agent evaluation into behavioral components like plan quality, adaptability, and inter-agent coordination. Another argues for longitudinal tracking, focusing on how agents evolve over time, whether they drift or stabilize, and whether they generalize or forget. If you are evaluating agents today, here are the most important criteria to measure: • 𝗧𝗮𝘀𝗸 𝘀𝘂𝗰𝗰𝗲𝘀𝘀: Did the agent complete the task, and was the outcome verifiable? • 𝗣𝗹𝗮𝗻 𝗾𝘂𝗮𝗹𝗶𝘁𝘆: Was the initial strategy reasonable and efficient? • 𝗔𝗱𝗮𝗽𝘁𝗮𝘁𝗶𝗼𝗻: Did the agent handle tool failures, retry intelligently, or escalate when needed? • 𝗠𝗲𝗺𝗼𝗿𝘆 𝘂𝘀𝗮𝗴𝗲: Was memory referenced meaningfully, or ignored? • 𝗖𝗼𝗼𝗿𝗱𝗶𝗻𝗮𝘁𝗶𝗼𝗻 (𝗳𝗼𝗿 𝗺𝘂𝗹𝘁𝗶-𝗮𝗴𝗲𝗻𝘁 𝘀𝘆𝘀𝘁𝗲𝗺𝘀): Did agents delegate, share information, and avoid redundancy? • 𝗦𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗼𝘃𝗲𝗿 𝘁𝗶𝗺𝗲: Did behavior remain consistent across runs or drift unpredictably? For adaptive agents or those in production, this becomes even more critical. Evaluation systems should be time-aware, tracking changes in behavior, error rates, and success patterns over time. Static accuracy alone will not explain why an agent performs well one day and fails the next. Structured evaluation is not just about dashboards. It is the foundation for improving agent design. Without clear signals, you cannot diagnose whether failure came from the LLM, the plan, the tool, or the orchestration logic. If your agents are planning, adapting, or coordinating across steps or roles, now is the time to move past simple correctness checks and build a robust, multi-dimensional evaluation framework. It is the only way to scale intelligent behavior with confidence.

  • View profile for Lauren McGoodwin

    Principle Content Strategist @ Atlassian | Brand & Content Marketing | AI Content Creator | Speaker & Author | Podcast Host

    31,108 followers

    I’ve been interviewing candidates for a new role and there’s one thing I’ve seen 90% of them struggle with: sharing the story of their career achievements. But don’t worry—I’ve got a simple hack that can help you overcome it: ✏️ Create a monthly ritual to review and document every significant work win, and turn each into a mini-case study. Documenting your wins regularly will save you HOURS when you prep for your next interview—plus it’s great fodder for: ⤷ your annual performance review ⤷ your 1x1s with your manager ⤷ your resume Here’s my 3-step process: 1️⃣ Weekly Check-in: Turn work ➡️ wins ⤷ Start a weekly habit of documenting your wins (grab my free template in the comments). ⤷ Block 30 minutes on your calendar every Friday to hold yourself accountable. ⤷ Ask yourself, “What did I accomplish this week that moved the needle?”   2️⃣ Monthly Recap: Turn wins ➡️ headlines ⤷ Identify 1–2 significant achievements and summarize them using this formula: [Action Verb] + [Specific Metric] + [Timeframe] + [Business Impact] ⤷ Make a bullet-point list (so you can stay organized and repurpose it for your resume later!) ⤷ Include dates and timelines for your own records—you’ll use them in step 3.   3️⃣ Quarterly Story-Building: Headlines ➡️ stories ⤷ Identify your top 3 quarterly wins. ⤷ Start a fresh document and map out each of those wins using the STAR method: ️ ⭐ Situation: What was the context? ️⭐ Task: What was your specific responsibility? ⭐ Action: What steps did you take? ⭐ Result: What measurable outcome did you achieve? ⤷ Ask AI to help you share that information as a story. Here’s the prompt I like to use: ✍ Can you help me turn this achievement into a story using the STAR framework for an upcoming interview for a [title here] role? Please keep it concise. [paste win]   Here’s what this looks like in action 👇 ⤷ Weekly win: March ’23 → Decreased CPA by 28% & increased conversion by 15% ⤷ Monthly recap: Optimized paid search campaigns in March 2023 that decreased CPA by 28% while increasing conversions by 15%, resulting in higher profit margins for the company. ⤷ Quarterly story: When I joined the marketing team in January 2023, our paid search campaigns were generating leads but at a high CPA, with budget constraints approaching in Q2.I was tasked with reducing CPA without sacrificing lead volume. In March 2023, I audited our campaigns and implemented three key changes: restructured ad groups with tightly-themed keywords, refined match types with strategic negative keywords, and A/B tested value-focused ad copy. By month-end, these optimizations decreased cost-per-acquisition by 28% while increasing conversion volume by 15%, saving budget and creating a scalable framework for future campaigns. What are your tips for storytelling in your interviews? I’d love to hear them. 

  • View profile for Roxanne Bras Petraeus
    Roxanne Bras Petraeus Roxanne Bras Petraeus is an Influencer

    CEO @ Ethena | Helping Fortune 500 companies build ethical & inclusive teams | Army vet & mom

    24,735 followers

    If you're embarking on a big initiative in 2025, be it professional or personal, I strongly recommend sending a monthly update (even if you're literally just emailing yourself). Here's the exact structure I've used for a few years: Context I send a monthly update to Ethena's investors. While I'm contractually obligated to do this, it's a phenomenal exercise because I'm forced to zoom out and assess progress. The monthly cadence is perfect because it's enough time for there to be something significant to say, but not so frequent that it becomes a burden. Structure 1. The TLDR/summary. No more than 3 bullet points summarizing what I think the biggest developments are. This is fuzzy and based totally on my intuition. 2. The metrics. These *have* to be the same metrics every month. I report on 8 key metrics and if I ever change a metric, I force myself to explain why I'm changing, say, how we calculate gross dollar retention. This builds accountability. 3. Team updates. It always sounds corny, but people will make or break your goal. While this is obviously true in business, I'd argue it's true even in personal goal setting. Want to get fit? You'll need to find the right coach. So I write what's going well (and not), and what open roles we have. 4. Biggest challenge. 2-3 sentences on whatever is hardest. 5. Asks. I ask my investors for help every single month. 6. Thanks. I thank everyone who did something in the past month. This is really important! It builds gratitude and people like being seen for their contributions. One last thing I do before I send an update is I read my previous month's. It helps me to see the through line and also, it's nice to see progress so concretely. I hope you read all the things/lift all the weights/accomplish whatever it is you're excited to tackle in 2025! And LMK if you think my update is missing something.

  • View profile for Pan Wu
    Pan Wu Pan Wu is an Influencer

    Senior Data Science Manager at Meta

    52,270 followers

    A "sampled success metric" is a performance measure or evaluation criterion calculated from a sample or subset of data rather than the entire population. Its calculation often involves higher costs per sample, such as manual review, leading to a trade-off between sample size and metric accuracy/sensitivity. In this tech blog, written by the data science team from Shopify, the discussion revolves around how the team leverages Monte Carlo simulation to understand metric variability under various scenarios to help the team make the right trade-offs. Initially, the team defines simulation metrics to describe the variability of the sampled success metric. For instance, if the actual success metric is decreasing over time, the metric could indicate how many months of sampled success metric would show a decrease, termed as "1-month decreases observed". Then, the team defines the distribution to run the Monte Carlo simulation. Monte Carlo simulation, a computational technique using random sampling to estimate outcomes of complex systems or processes with uncertain inputs, draws samples from a dedicated distribution that matches business needs. Based on past observations, the team’s application follows a Poisson distribution. Next comes the massive simulation phase, where the team runs multiple simulations for one parameter and then changes various parameters to simulate different scenarios. The goal is to quantify how much the sample mean will differ from the underlying population mean given realistic assumptions. The final result provides a clear statistical distribution of how much extra sample size could lead to metrics variability decrease and increased accuracy. This case study demonstrates that Monte Carlo simulation could be a valuable toolkit to add to your decision-making and data science knowledge. #datascience #analytics #metrics #algorithms #simulation #montecarlo #decisionmaking – – –  Check out the "Snacks Weekly on Data Science" podcast and subscribe, where I explain in more detail the concepts discussed in this and future posts:    -- Spotify: https://lnkd.in/gKgaMvbh   -- Apple Podcast: https://lnkd.in/gj6aPBBY    -- Youtube: https://lnkd.in/gcwPeBmR https://lnkd.in/dKnrZzzV 

  • View profile for Tyler Folkman
    Tyler Folkman Tyler Folkman is an Influencer

    Chief AI Officer at JobNimbus | Building AI that solves real problems | 10+ years scaling AI products

    19,196 followers

    I audited 464 of my AI agent sessions. The numbers were worse than I expected. 1.2B tokens. $255.73 in spend. 26% subagent failure rate. 97% of sessions never compacted. One session had 39 consecutive tool errors. One broken command retried 12 times with no stdin. The embarrassing part: all of this data was already sitting in my logs. I had spent months improving prompts, trying models, tweaking configs, and building workflows. But I had not built the habit that actually matters once agents become part of daily work: Read the system. Not vibes. Not demos. Not "this felt faster." Actual logs: 1. Where did the agent spend tokens? 2. Which tools failed repeatedly? 3. Which subagents completed useful work? 4. Where did context management break? 5. Which sessions should have stopped earlier? This is why I think harness engineering is becoming the successor to prompt engineering. The bottleneck is no longer "can the model do the task?" The bottleneck is whether the system around the model creates visibility, constraints, feedback loops, and stop conditions. I wrote up the full audit here: https://lnkd.in/gTn9WA-y If your team uses AI coding agents, what are you actually measuring: model quality, or the workflow around the model?

  • View profile for Greg Coquillo

    AI Platform & Infrastructure Product Leader | Scaling massive AI Factories for Frontier Model providers | Azure AI & HPC | Former AWS, Amazon | Startup Investor | I deploy GPU-as-a-Service for AI customers

    234,344 followers

    ✋Before rushing into training models, do not skip the part that actually determines whether the model is useful: Measuring performance. Without the right metrics you are not evaluating a model, you are just validating your assumptions. Check out theses nine metrics every ML practitioner should understand and use with intention 👇 1. Accuracy Good for balanced datasets. Misleading when classes are skewed. 2. Precision Of the samples you predicted as positive, how many were correct. Important when false positives are costly. 3. Recall Of the samples that were actually positive, how many you caught. Critical when false negatives are dangerous. 4. F1 Score Balances precision and recall. Reliable when you need a single metric that reflects both types of error. 5. ROC AUC Measures how well a model separates classes across thresholds. Useful for model comparison independent of cutoffs. 6. Confusion Matrix Exposes the exact distribution of true positives, false positives, true negatives, and false negatives. Great for diagnosing failure modes. 7. Log Loss Penalizes confident wrong predictions. Important for probabilistic models where calibration matters. 8. MAE (Mean Absolute Error) Average of absolute errors. Simple, interpretable, and robust for many regression problems. 9. RMSE (Root Mean Squared Error) Heavily penalizes large errors. Best when you care about avoiding big misses. Strong ML systems are built by measuring the right things. These metrics show you how your model behaves, where it fails, and whether it is ready for production. What else would you add? #AI #ML

  • View profile for Jack Appleby
    Jack Appleby Jack Appleby is an Influencer

    Social Media / Creator Consultant | Work: Microsoft, Beats By Dre, Verizon, Twitch, Morning Brew, Rock Band, Community (six seasons and a movie!) & a slew of video games.

    89,204 followers

    The first 2 months of 2025 are over, so every professional needs to start recapping their year for LinkedIn / their resume (seriously). Here's how: Write down your biggest professional accomplishments for January and February. You want to write 1 sentence each for: - What the project was - Your role in the project - Results You really should be doing this at the end of every major work project, at the end of each month, or whenever you feel like you accomplished something! Writing the immediate 3 sentence case study lets you cement the details longggg before you need them for some resume 2 years down the line. Now, where do you keep them? Great, glad you asked. You could either: - gather a google doc that you keep updated - just text them to yourself so they're in your phone - or, if the project is one of the four best things you've done at your current job, add it to your job description on LinkedIn! (if you want an example, go to my profile and look how I did the Twitch years) Get in this habit of grabbing your own achievements as they happen. It's great for documentation for resumes, raise and promotion negotiations, and just good ol' self-gratitude.

  • View profile for Sriram Natarajan

    Engineering at Google - Gemini Enterprise, TEDx Speaker

    4,074 followers

    𝗘𝘃𝗮𝗹 𝗶𝘀𝗻’𝘁 𝗤𝗔. 𝗜𝘁’𝘀 𝘁𝗵𝗲 𝗻𝗲𝘄 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗔𝗜 𝘀𝗸𝗶𝗹𝗹𝘀𝗲𝘁. Last week, OpenAI rolled back a GPT-4o update because it got… too agreeable. The model started endorsing user biases just to keep people happy. They called it “𝘀𝘆𝗰𝗼𝗽𝗵𝗮𝗻𝗰𝘆.” 🔗 Checkout: https://lnkd.in/gws7tBRe 𝗪𝗵𝘆 𝗱𝗶𝗱 𝘁𝗵𝗶𝘀 𝗵𝗮𝗽𝗽𝗲𝗻? They over-indexed on thumbs-up feedback, not deeper evaluations. And it broke alignment. This is the iceberg enterprise AI teams are quietly sailing toward. AI systems don’t just fail because of hallucinations. They fail because apps don’t test what actually matters, 𝙘𝙤𝙣𝙩𝙞𝙣𝙪𝙤𝙪𝙨𝙡𝙮. 𝗪𝗵𝗮𝘁 𝗰𝗮𝗻 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲𝘀 𝗹𝗲𝗮𝗿𝗻? Most teams treat evaluation like a test gate. But LLM-integrated systems evolve: ✔️ Model behavior drifts ✔️ User inputs are unpredictable ✔️ Success criteria shift with every rollout And here’s the nuance many miss: 𝗕𝗹𝗮𝗻𝗸𝗲𝘁 𝗣𝗨𝗡𝗧𝘀 𝗵𝗶𝗱𝗲 𝗺𝗼𝗱𝗲𝗹 𝘃𝗮𝗹𝘂𝗲. 𝗟𝗼𝗼𝘀𝗲𝗻 𝗳𝗶𝗹𝘁𝗲𝗿𝘀, 𝗮𝗻𝗱 𝘆𝗼𝘂 𝗿𝗶𝘀𝗸 𝘁𝗿𝘂𝘀𝘁. 👉 The fix: Precision in evals is a good way to scale safely. 𝗘𝘃𝗮𝗹 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘀: The missing enterprise skillset Think: 🧠 Solution Architect + 🎯 QA Strategist + 🧩 Product Thinker — but for AI systems. This isn’t QA. It’s a 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝗶𝗰 𝗰𝗮𝗽𝗮𝗯𝗶𝗹𝗶𝘁𝘆. Here’s what Eval Architects can enable: ✅ 𝗗𝗲𝗳𝗶𝗻𝗲 𝗔𝗡𝗗 𝗲𝘃𝗼𝗹𝘃𝗲 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗺𝗲𝘁𝗿𝗶𝗰𝘀 Draw from Anthropic's Claude success criteria (https://lnkd.in/gaS9ctXb) and OpenAI's Preparedness Framework (https://lnkd.in/grrYpr5v) — don’t just define KPIs once. 𝘽𝙪𝙞𝙡𝙙 𝙛𝙤𝙧 𝙙𝙧𝙞𝙛𝙩. ✅ 𝗘𝘃𝗼𝗹𝘃𝗲 𝘁𝗲𝘀𝘁𝗶𝗻𝗴 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝘆 𝘁𝗼 𝗿𝘂𝗻 𝗻𝗶𝗴𝗵𝘁𝗹𝘆 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲𝘀 Include tests for reasoning, features, product outcomes, and business impact. ✅ 𝗗𝗲𝘀𝗶𝗴𝗻 𝗮𝗻𝗱 𝗱𝗲𝗽𝗹𝗼𝘆 𝗿𝗲𝗮𝗹-𝘁𝗶𝗺𝗲 𝗲𝘃𝗮𝗹 𝗮𝗴𝗲𝗻𝘁𝘀 Not just to monitor — but to 𝘭𝘦𝘢𝘳𝘯 from anonymized production data, identify failure modes, and suggest new test coverage. This is how evaluation becomes 𝗮𝗱𝗮𝗽𝘁𝗶𝘃𝗲, not reactive. 𝗦𝗼𝗹𝘂𝘁𝗶𝗼𝗻 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘀 and 𝗣𝗿𝗼𝗱𝘂𝗰𝘁 𝗢𝗽𝘀 leaders are well-positioned to grow into this. They already think in systems, metrics, and risk. Now they need to think in 𝗲𝘃𝗮𝗹𝘀.

Explore categories