Science Forecasting Models

Explore top LinkedIn content from expert professionals.

  • View profile for Pan Wu
    Pan Wu Pan Wu is an Influencer

    Senior Data Science Manager at Meta

    52,270 followers

    In today’s fast-paced tech landscape, understanding the true causal impact of business decisions is more critical than ever. Whether you're launching a new feature, running a marketing campaign, or testing operational changes, it’s essential to go beyond correlation and uncover what actually drives outcomes. In a recent blog post, a data scientist from Walmart explains what Bayesian Structural Time Series (BSTS) models are and how they can be used to measure causal impact. BSTS is a flexible time series modeling approach that breaks down data into components like trend, seasonality, and regressors—enabling teams to simulate what would have happened without an intervention. The post does a great job of explaining the methodology with clear, real-world examples. It’s a valuable read for anyone working on experimentation, marketing measurement, or causal inference at scale. #DataScience #MachineLearning #CausalInference #Analytics #BayesianModeling #SnacksWeeklyonDataScience – – –  Check out the "Snacks Weekly on Data Science" podcast and subscribe, where I explain in more detail the concepts discussed in this and future posts:    -- Spotify: https://lnkd.in/gKgaMvbh   -- Apple Podcast: https://lnkd.in/gFYvfB8V    -- Youtube: https://lnkd.in/gcwPeBmR https://lnkd.in/gzSZcSh8

  • View profile for Corrado Botta

    Postdoctoral Researcher

    13,759 followers

    BAYESIAN REGRESSION: FROM POINT ESTIMATES TO PROBABILITY DISTRIBUTIONS 📊 In empirical finance and economics, classical OLS regression often gives us false confidence through point estimates that ignore parameter uncertainty. When sample sizes are small or data is noisy—common in emerging markets or early-stage ventures—this overconfidence can lead to costly decisions. 🎯 The Bayesian approach transforms regression from a single "best fit" line into a rich distribution of plausible relationships, naturally quantifying our uncertainty about model parameters. The fundamental shift in thinking: Classical: "The true slope is 1.47 (±0.23)" Bayesian: "Given the data, we believe the slope is most likely around 1.47, with 90% probability between 1.05 and 1.89" This probabilistic framework offers three key advantages for applied research: 📈 Prior Integration: Incorporate domain expertise or previous studies directly into the analysis—invaluable when working with limited data or combining multiple information sources 🔄 Natural Uncertainty Propagation: Parameter uncertainty flows seamlessly into predictions, giving honest confidence intervals that reflect both estimation uncertainty and inherent variability 📊 Richer Inference: Extract any quantity of interest from posterior distributions—tail risks, probability of economic significance, or decision-theoretic optimal choices Grid approximation, while computationally limited to low dimensions, provides profound value. By discretizing the parameter space and computing posteriors explicitly, we demystify the "black box" of Bayesian inference—making it accessible to practitioners and stakeholders alike. Real-world applications where this matters: • Estimating risk premia with short time series • Policy evaluation with limited pilot data • Cross-border investment decisions under regime uncertainty • Incorporating expert judgment in forensic economics • Robust forecasting when historical relationships may be shifting The beauty lies not in abandoning classical methods, but in acknowledging when uncertainty quantification becomes as important as point estimation itself. Currently exploring applications in financial econometrics and decision science—always interested in connecting with researchers and practitioners tackling similar challenges! What domains in your work could benefit from honest uncertainty quantification? 🤔 #BayesianEconometrics #QuantitativeFinance #DataScience #RiskAnalysis #EmpiricalResearch #StatisticalModeling

  • View profile for Dr. Olivera Stojanovic

    Founder @ spacebayes | Bayesian Modeling for Earth Observation | Climatebase Fellow

    2,170 followers

    𝗪𝗵𝘆 𝗕𝗮𝘆𝗲𝘀𝗶𝗮𝗻 𝗛𝗶𝗲𝗿𝗮𝗿𝗰𝗵𝗶𝗰𝗮𝗹 𝗠𝗼𝗱𝗲𝗹𝘀 𝗔𝗿𝗲 𝗦𝗼 𝗣𝗼𝘄𝗲𝗿𝗳𝘂𝗹 A common challenge in data science is dealing with #heterogeneous data, because different regions, customer segments, or product categories may have vastly different amounts of data. Traditional approaches either 𝗺𝗼𝗱𝗲𝗹 𝗲𝗮𝗰𝗵 𝗴𝗿𝗼𝘂𝗽 𝘀𝗲𝗽𝗮𝗿𝗮𝘁𝗲𝗹𝘆, leading to noisy estimates when data is scarce, or force a 𝘀𝗶𝗻𝗴𝗹𝗲 𝗺𝗼𝗱𝗲𝗹 𝗮𝗰𝗿𝗼𝘀𝘀 𝗮𝗹𝗹 𝗴𝗿𝗼𝘂𝗽𝘀, ignoring real differences. 𝗕𝗮𝘆𝗲𝘀𝗶𝗮𝗻 𝗵𝗶𝗲𝗿𝗮𝗿𝗰𝗵𝗶𝗰𝗮𝗹 𝗺𝗼𝗱𝗲𝗹𝘀 offer a different solution. They allow parameters to vary at 𝗺𝘂𝗹𝘁𝗶𝗽𝗹𝗲 𝗹𝗲𝘃𝗲𝗹𝘀 𝗼𝗳 𝗱𝗮𝘁𝗮 𝗿𝗲𝗹𝗮𝘁𝗶𝗼𝗻𝘀𝗵𝗶𝗽𝘀, letting us incorporate not just the data itself but also its underlying structure, #metadata, and the way it was collected. They capture shared #patterns while accounting for group-specific differences. This flexibility makes them ideal for data that’s nested or structured across multiple dimensions. In 𝗲𝗻𝘃𝗶𝗿𝗼𝗻𝗺𝗲𝗻𝘁𝗮𝗹 𝘀𝗰𝗶𝗲𝗻𝗰𝗲, Bayesian hierarchical models are widely used because they allow scientists to measure effects at different locations, over time, or at different latitudes, all while capturing broader trends. You can read about such one example here: https://lnkd.in/d6ERwa7q In a business use case, such as 𝗿𝗲𝘁𝗮𝗶𝗹 𝗱𝗲𝗺𝗮𝗻𝗱 𝗳𝗼𝗿𝗲𝗰𝗮𝘀𝘁𝗶𝗻𝗴, Bayesian hierarchical models provide: • 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻 𝗼𝗳 𝗱𝗮𝘁𝗮 𝗮𝗰𝗿𝗼𝘀𝘀 𝗿𝗲𝗴𝗶𝗼𝗻𝘀, 𝘀𝘁𝗼𝗿𝗲𝘀, 𝗮𝗻𝗱 𝗽𝗿𝗼𝗱𝘂𝗰𝘁 𝗰𝗮𝘁𝗲𝗴𝗼𝗿𝗶𝗲𝘀, capturing both global trends and local variations. • 𝗦𝗲𝗮𝘀𝗼𝗻𝗮𝗹𝗶𝘁𝘆 𝗺𝗼𝗱𝗲𝗹𝗶𝗻𝗴, assuming common patterns across regions but also allowing for regional differences. • 𝗛𝗮𝗻𝗱𝗹𝗶𝗻𝗴 #sparse 𝗱𝗮𝘁𝗮, borrowing information from related datasets to improve #accuracy. You can read more about this application: https://lnkd.in/dnkcKi4b In both cases, I used #PyMC for Bayesian modeling. By allowing flexibility and borrowing strength from related data, Bayesian hierarchical models offer a robust approach to #forecasting, 𝗲𝘃𝗲𝗻 𝘄𝗶𝘁𝗵 𝗹𝗶𝗺𝗶𝘁𝗲𝗱 𝗼𝗿 𝘂𝗻𝗲𝘃𝗲𝗻 𝗱𝗮𝘁𝗮. Let me know if you've used Bayesian hierarchical models, I'd love to hear about other use cases. #BayesianInference #HierarchicalModels #DataScience #MachineLearning #Forecasting #RetailAnalytics #PyMC #EnvironmentalScience #DemandForecasting #StatisticalModeling #BusinessAnalytics #GeospatialModeling #PredictiveModeling #DataAnalysis 

  • View profile for Dr. Juan Camilo Orduz

    Applied Scientist | Ph.D. Math | Open Source

    9,644 followers

    Here is the pre-print of the Bayesian approach to cohort revenue and retention modeling 🙂 https://lnkd.in/dtgAywWD “We present a Bayesian approach to model cohort-level retention rates and revenue over time. We use Bayesian additive regression trees (BART) to model the retention component which we couple with a linear model for the revenue component. This method is flexible enough to allow adding additional covariates to both model components. This Bayesian framework allows us to quantify uncertainty in the estimation, understand the effect of covariates on retention through partial dependence plots (PDP) and individual conditional expectation (ICE) plots, and most importantly, forecast future revenue and retention rates with well-calibrated uncertainty through highest density intervals. We also provide alternative approaches to model the retention component using neural networks and inference through stochastic variational inference.” #cohort #retention #clv #bayes

  • View profile for Nikhil Dhand

    Creator of Probabilistic Chain Analysis | Helped teams decode why megaprojects really fail | 16,000+ projects across 8 sectors | Peer reviewed Author | PMP | WSP

    6,219 followers

    I just published a research paper that challenges how we model risk. And the result will make most project managers uncomfortable. ↓ Standard Monte Carlo assumes risks fire independently. They don't. They fire in chains. Risk A delays procurement. Procurement delay pushes mobilisation. Mobilisation delay compresses testing. Compressed testing forces rework. Rework blows contingency. That's not bad luck. That's a cascade. And your risk register cannot see it. ━━━━━━━━━━━━━━━━ I spent months building a framework to model exactly this — Probabilistic Chain Analysis (PCA). The result from a UK highways case study: ▸ One pre-mobilisation intervention ▸ £250,000 reduction in P90 cost exposure ▸ £15,000 management cost ▸ 16.7x return on risk management effort Not because we worked harder. Because we looked at the right node. ━━━━━━━━━━━━━━━━ The methodology: → Map risks as a Directed Acyclic Graph (not a flat register) → Assign conditional probabilities using Bayesian Networks → Run coupled Monte Carlo simulation → Identify cascade lift factor — which node is amplifying everything? → Intervene there. Not everywhere. ━━━━━━━━━━━━━━━━ The paper is 27 pages. Open access. No paywall. Validated across 16,000 infrastructure projects across 8 sectors. UK highways. Solar EPC. Hospitals. Power plants. Same pattern every time. The most dangerous risk isn't the most probable one. It's the most connected one. ━━━━━━━━━━━━━━━━ 📄 Full paper (free): in comments ━━━━━━━━━━━━━━━━ Have you ever seen a cascade take down a project that looked fine on paper? Drop it in the comments. I read every one. #ProjectManagement #RiskManagement #MonteCarlo #BayesianNetworks #Infrastructure #EPC #ProjectControls #PMP #Quantitative

  • View profile for Alexander Denev

    CEO & Founder | Turnleaf Analytics | Macro Forecasting Infrastructure | Institutional Quant

    9,211 followers

    When modelling the impact of events that have never happened before, historical data fails us. No time series can tell you what happens when the Strait of Hormuz — through which 20% of the world's oil flows — is effectively shut for the first time in history. This is exactly the kind of problem Bayesian Networks were designed for. Bayesian Networks (BNs) allow us to aggregate forward-looking probabilities into a coherent whole by decomposing a complex scenario into smaller, interacting parts — each informed by whatever evidence we have: expert judgement, market data, prediction markets, economic theory. The language of probability keeps the picture internally consistent. The graphical structure makes the assumptions visible and debatable. And the output — a joint probability distribution — can be fed directly into portfolio optimisation, stress testing, or hedging decisions. Here I show a practical example, built today. Starting from the Iran-US ceasefire negotiations as the root node, I trace the cascade of consequences through the Strait of Hormuz blockade, oil prices, US inflation, Fed policy, GDP growth, and finally to the four risk factors that matter for a US portfolio: equities, Treasuries, credit spreads, and the dollar. The conditional probability tables were populated using information gathered in real time from the web — news sources, IMF forecasts, prediction markets, analyst reports — consolidated in minutes with the help of Claude. What would have taken days of manual research and spreadsheet wiring now takes a conversation. The resulting unconditional marginals tell a clear story: all four terminal risk factors are skewed to the adverse side. The modal path through the network points to a mild stagflation scenario — oil at $80–100, inflation at 3–4%, the Fed trapped on hold, equities in sell-off territory, yields rising, credit spreads widening. But the tails are where it gets interesting — and where BNs earn their keep. The probability of an equity crash exceeds 15%. The probability of oil above $130 is around 20%. These are the fat tails that no VaR model calibrated on pre-war data would capture. The approach is not a crystal ball. The probabilities are approximate. The structure reflects judgement calls. But that is precisely the point: Bayesian Networks make every assumption explicit, every judgement auditable, and every sensitivity testable.

  • View profile for Brenden Delarua

    Chief Marketing Officer @ Stella Incrementality & MMM

    11,515 followers

    You can’t just do “test minus control” and call it incrementality. That’s not a model. That’s a middle school math problem. Or at best, it’s an entry level difference-in-differences analysis. Yes, technically, comparing a test group to a control group can work, if you’ve perfectly matched them and nothing weird happens. No weather issues, no influencer shoutouts, no backend outages. No noise. But that’s rarely the case. Real life is messy. And models exist to clean that mess up. BSTS (Bayesian Structural Time Series) doesn’t just subtract. It builds a counterfactual. A “what would have happened” scenario based on trends, seasonality, day-of-week patterns, and a bunch of invisible math gremlins doing some very serious forecasting. It says, “Based on the last 4 months of data and what’s happening in similar markets right now, here’s what revenue should’ve looked like in the test group.” Then it compares actuals to that forecast, complete with error bands so you’re not just guessing. It’s not perfect, but it’s a hell of a lot more reliable than eyeballing a revenue chart and calling it science. Another approach we use is weighted synthetic control. Instead of choosing a single control market and hoping it behaves, you build a custom control made up of multiple regions, each weighted based on how well they historically match the test group. So instead of saying “Denver is our control,” you’re saying “It’s 20 percent Denver, 30 percent Austin, 10 percent Boise,” and so on. Different methods, same goal: a cleaner, more accurate baseline. In other words, we’re not just subtracting numbers. We’re accounting for all external factors that could give us false positives. Because when millions in ad spend are on the line, you don’t just want an answer. You want an accurate one. And “test minus control” might look simple, but unless you’ve controlled for everything else, it’s usually wrong. Tools like Stella make this incredibly simple and surprisingly affordable for marketing teams. Message me if you want more accurate marketing measurement.

  • View profile for Massoud Amin

    Envisioned, funded, and led the R&D behind the smart self-healing grid — 28 years and building | Writing on what holds | Author, Both Your Houses | Professor Emeritus, University of Minnesota | IEEE & ASME Fellow

    11,858 followers

    Forecasting Risk in Today’s Power System Electricity prices follow human decisions, not formulas. They move with the weather, demand, fuel cost, and strategy. In 2012, we built a model that used Bayesian learning and stochastic games to forecast price distributions rather than single points. It worked then. It’s essential now. The system has changed. North America’s grid is managed through six NERC regional entities. ISOs and RTOs run about two-thirds of U.S. demand. Market operators now rely on probabilistic and Monte Carlo analysis for planning, pricing, and reliability. The old deterministic view is gone. The numbers show the shift. U.S. electricity demand set records in 2024 and again in 2025. Growth comes from data centers, electric vehicles, and manufacturing. The U.S. will add 63 gigawatts of new capacity this year, 81 percent of which will come from solar and batteries. Utility-scale storage will pass 65 GW by 2026. Renewables’ share of generation will climb from 23 percent in 2024 to 27 percent in 2026. Natural gas will decline toward 39 percent, and coal will fall below 14 percent. The key lessons remain. 1. Learn continuously. Bayesian updating incorporates new data—weather, bids, outages—to keep forecasts up to date. 2. Model real behavior. Prices form from competing decisions under limits, not from ideal equations. 3. Show the full range. A probability curve gives investors, traders, and planners the truth about exposure and resilience. The tools are better. GPU computing and scenario reduction now make real-time probabilistic forecasting routine. ISOs use stochastic unit commitment and risk-based adequacy methods. These drive real investment and operational choices, not academic models. The outcome is clear. Forecasting means measuring uncertainty, not hiding it. The most resilient organizations are those that see risk early, price it correctly, and act before others react. We forecast risk because risk drives every real decision—capital, reliability, and trust. The grid’s future will belong to those who treat uncertainty as information, not noise. — Sources: NERC State of Reliability 2025; EIA Today in Energy (May–Oct 2025); FERC Market Reports; ISO/RTO Council Data; Amin & Peck Probabilistic Price Model. #AI #Analytics #Bayesian #Data #Energy #Engineering #Foresight #Forecasting #Grid #Innovation #Leadership #Mathematics #Modeling #Optimization #Probability #Resilience #Risk #Simulation #Sustainability #Systems #Technology

  • View profile for Geetha Malika N

    Epidemiologist | RWE Data Scientist | HEOR | Causal Inference | Medical Devices | FDA/CDRH

    7,321 followers

    ✍️ Day 4/30 Continuation from Day 3... When we borrow from historical trial data, there are two main approaches: Frequentist and Bayesian. 🔹 Frequentist Borrowing = Bring the old patients in Here, you literally reuse patient-level data from earlier trials as if those patients were enrolled in your new trial. Main methods: 1️⃣ Pooled Analysis All patients from past and present trials are combined into one dataset. This increases sample size and power, but works best when trial designs and populations are very similar. 2️⃣ External / Synthetic Controls When a new trial doesn’t have a control arm (common in oncology or rare diseases), you create one using data from past RCTs, registries, or real-world databases. This gives you a comparator without running a new randomized control. 3️⃣ Test-Then-Pool First, you statistically test whether the historical control group looks like the new control group. If they match, you pool them; if not, you keep them separate. This ensures you only borrow when history aligns with the present. 🔹 Bayesian Borrowing = Bring the old evidence in Instead of reusing patients, you reuse the knowledge from past trials, their means, variances, and distributions, which form a prior. The new trial provides a likelihood, and together they update into a posterior (our new best estimate). Main methods: 1️⃣ Power Priors You give historical data “partial credit” by assigning a weight (α). α = 1 means fully trust history; α = 0 means ignore it; anything in between is partial borrowing. 2️⃣ MAP Priors (Meta-Analytic Predictive) When you have multiple past studies, you combine them via meta-analysis to create a single predictive prior. This reflects the average evidence across all past trials. 3️⃣ Commensurate Priors These adapt the amount of borrowing depending on how similar past and present data look. If the two agree, you borrow more; if not, you borrow less. 4️⃣ Hierarchical Models You treat each study as related but not identical. The model learns how much to borrow for each dataset automatically, more from similar studies, less from outliers. 5️⃣ Robust Priors (e.g., Robust MAP) You blend an informative prior (from history) with a weak or skeptical prior. This way, if the historical evidence is misleading, the weaker prior prevents over-borrowing. 💡 High-level picture: Frequentist = patients → direct, transparent, but only works well when trials are comparable. Bayesian = evidence → flexible, adaptive, and can handle more complex scenarios, but requires careful modeling. 👉 Tomorrow, I’ll dive deeper into the Frequentist methods, with real-world trial examples from PubMed.

  • 📊 #Bayesian Hierarchical Models for Multi-Market Marketing Measurement In global #marketinganalytics, one of the biggest challenges is measuring performance across markets with vastly different scales, budgets, and consumer behavior - without losing #statisticalpower. That’s where #BayesianHierarchicalModels (BHMs) come in. Instead of analyzing each market independently or pooling all data together, BHMs share information across markets in a probabilistic way — letting data-rich markets “inform” data-sparse ones while preserving local differences. 🔍 Why it matters Traditional #regression-based Marketing Mix Models (MMMs) often treat markets separately, leading to unstable or noisy ROI estimates - especially for small markets. 📈 Example Imagine 20 countries running similar paid media campaigns. A Bayesian hierarchical MMM can: - Estimate global channel effectiveness (e.g., search, display, video) - Allow market-level deviations (e.g., paid search works 2× better in Japan) - Quantify uncertainty per market - so decision-makers can trust confidence intervals, not point estimates. 💥 Business impact - Enables cross-market ROI benchmarking - Reduces overfitting and noise in low-volume markets - Supports global budget optimization with credible intervals - Provides a unified model for decision-making rather than fragmented reports 🧠 Takeaway Bayesian Hierarchical Modeling turns scattered marketing data into a coherent global learning system. It’s the bridge between local insights and global strategy — enabling marketers to scale spend decisions with statistical confidence. #MarketingAnalytics #BayesianModeling #MMM #DataScience #MarketingOptimization #CausalInference #AnalyticsLeadership

Explore categories