Best LLM-based Open-Source tool for Data Visualization, non-tech friendly CanvasXpress is a JavaScript library with built-in LLM and copilot features. This means users can chat with the LLM directly, with no code needed. It also works from visualizations in a web page, R, or Python. It’s funny how I came across this tool first and only later realized it was built by someone I know—Isaac Neuhaus. I called Isaac, of course: This tool was originally built internally for the company he works for and designed to analyze genomics and research data, which requires the tool to meet high-level reliability and accuracy. ➡️Link https://lnkd.in/gk5y_h7W As an open-source tool, it's very powerful and worth exploring. Here are some of its features that stand out the most to me: 𝐀𝐮𝐭𝐨𝐦𝐚𝐭𝐢𝐜 𝐆𝐫𝐚𝐩𝐡 𝐋𝐢𝐧𝐤𝐢𝐧𝐠: Visualizations on the same page are automatically connected. Selecting data points in one graph highlights them in other graphs. No extra code is needed. 𝐏𝐨𝐰𝐞𝐫𝐟𝐮𝐥 𝐓𝐨𝐨𝐥𝐬 𝐟𝐨𝐫 𝐂𝐮𝐬𝐭𝐨𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧: - Filtering data like in Spotfire. - An interactive data table for exploring datasets. - A detailed customizer designed for end users. 𝐀𝐝𝐯𝐚𝐧𝐜𝐞𝐝 𝐀𝐮𝐝𝐢𝐭 𝐓𝐫𝐚𝐢𝐥: Tracks every customization and keeps a detailed record. (This feature stands out compared to other open-source tools that I've tried.) ➡️Explore it here: https://lnkd.in/gk5y_h7W Isaac's team has also published this tool in a peer-reviewed journal and is working on publishing its LLM capabilities. #datascience #datavisualization #programming #datanalysis #opensource
Data Analysis In Biology
Explore top LinkedIn content from expert professionals.
-
-
In global health, one number can influence millions of lives. But only when we interpret it correctly. One of the most misunderstood yet most powerful concepts in epidemiology and biostatistics is the Odds Ratio (OR). From outbreak investigations and vaccine effectiveness studies to nutrition surveillance, maternal health, HIV programs, and humanitarian response evaluations, Odds Ratios help us answer a critical question: How strongly is an exposure associated with an outcome? The challenge is that many professionals memorize formulas without fully understanding the real-world interpretation behind them. But in public health, interpretation is everything. That is why I developed this visual guide on Odds Ratio simplifying: 1. What OR actually means 2. Step-by-step calculations 3. Interpretation of OR < 1, OR = 1, and OR > 1 4. Case-control study applications 5. Public health examples and scenarios 6. Common interpretation mistakes 7. Epidemiological importance in surveillance and decision-making In disease surveillance and program evaluation, statistics should never remain theoretical. They must translate into: Better targeting of vulnerable populations Smarter prevention strategies Stronger evidence-based policies Faster outbreak response More efficient resource allocation Data alone does not save lives. Correct interpretation of data does. As epidemiology increasingly intersects with AI, data science, humanitarian programming, and global development, statistical literacy is no longer optional it is a leadership competency. I would be interested to hear from professionals across: Public Health | Epidemiology | Monitoring & Evaluation | Data Science | Clinical Research | Humanitarian Response | Global Health Policy In your experience, what is the most commonly misunderstood statistical concept in public health practice? #Epidemiology #PublicHealth #GlobalHealth #Biostatistics #DataScience #MonitoringAndEvaluation #HealthSystems #DiseaseSurveillance #ResearchMethods #EvidenceBasedPolicy #HealthEquity #HumanitarianResponse #UNICEF #WHO #SDGs #GlobalDevelopment #Statistics #OddsRatio #FieldEpidemiology
-
I spent the last few days rebuilding 20 of the figure types frequently appear in Nature and Cell papers, all in R with ggplot2. Manhattan and volcano plots, Circos diagrams, Sankey flows, raincloud and split-violin plots, treemaps, Mantel heatmaps, and more. Each figure comes with its own data simulator, so the whole thing runs on a fresh clone with nothing to download. Change one parameter and the figure updates with it. Drop in your own data when you're ready. I built it to get comfortable with publication-quality plotting for my genomics work, and to save the next person some of the trial and error. Github repo link for code and all 20 examples: https://lnkd.in/gHvgbr9X Acknowledgement: The idea for these figures came from a WeChat post shared by Dr. Rana Muhammad Atif. I then built the whole thing my own way, as a clean and freely accessible repository. #RStats #ggplot2 #Bioinformatics #DataVisualization #PlantScience #DataScience #Genomics #RStats #ComputationalBiology #PlantScience #Rprogramming
-
How to turn clinical data into real insight in your clinical evaluation? Let’s walk through it step by step.↴ 1. What is “clinical data”? Clinical data refers to any information regarding the safety or performance of a medical device that comes from: → Clinical investigations of the device itself → Other investigations or studies found in scientific literature → Post-market surveillance, especially PMCF → Peer-reviewed literature about similar clinical experience It’s not just “data”, it’s data that shows how the device behaves in the real world, in real hands, with real patients. 2. Clinical data must be relevant Relevance is contextual. That means: relevant to the medical device and its intended purpose. So: → Always match your data to a specific device and a specific clinical question → Exclude nonhuman, vague or off-label data unless justified → Avoid just dumping SOTA or competitor data unless it supports your argument → Carefully define search terms and clearly identify the device under discussion 3. Clinical data must be high quality Not all data is equal. High-quality clinical data should be: → Scientifically rigorous → Complete and methodologically sound → Controlled and peer-reviewed Use validated frameworks like: ✓ PICO (Population, Intervention, Comparator, Outcome) ✓Cochrane Handbook ✓ PRISMA ✓ MOOSE You can find guidance regarding appraisal of data in IMDRF, MDCG 2020-6, and Meddev 2.7/1. Sources from PubMed, Google Scholar, or Cochrane are preferred, especially if they’re top-tier journals. 4. Clinical data must be sufficient Here’s where many stumble. → Small sample size? Not enough → Low-quality studies or incomplete data? Weak evidence → Case reports and posters? Not generalizable → Not on EU population? Risky Sufficient data means: → Enough for an expert to form an opinion → Strong enough for the device’s risk level → Justified by hierarchy of evidence High-risk = needs deep literature and real-world experience Low-risk = might rely on limited, lower-quality data 5. So… when is clinical data enough? When it supports a clear, justified, and ethical approach to the device’s clinical evaluation. Risk-based is the key: → Higher risk = more data → Lower risk = less, if PMS and RM show acceptable benefit-risk
-
Bioinformatics Visualization in R. Part 2. Published Dataset Analysis: Immune Cell RNA Kinetics This second post covers reproduction of 7 adapted figures from the Xiong et al. (2025) article analysing single-cell RNA dynamics during acute enteric Salmonella infection. Scientists used new single-cell RNA labelling sequencing method, scIVNL-seq, to detect RNA synthesis and degradation interaction. This allowed to track RNA dynamics and control strategies in multiple immune cells across 72 hours of infection revealing temporal coordination of the immune response. Article: Xiong et al. scIVNL-seq resolves in vivo single-cell RNA dynamics of immune cells during Salmonella infection. Nat Commun. 2025; 16:7937. Seven visualisation types, one biological story. The challenge wasn't just to reproduce individual plots, but to understand how each panel answers a different question about immune response dynamics: - Boxplot: Which cell types are transcriptionally active? - Quadrant Scatter Plot: What RNA stability regimes exist? - Complex Heatmap: How do genes respond over time? - Pathway Heatmap: Which biological pathways activate when? - Bubble Plot: How does process complexity evolve? - Stacked Bar Chart: When does B to Plasma cells differentiation occur? - Network Graph: How do immune cells coordinate? Key technical decisions that matter. 1) Bubble Plot: Custom size legend The bubble plot needed a custom size legend (10, 20, 30 genes) that ggplot2 doesn't generate automatically. I applied manual construction using annotate(). Each reference bubble requires annotate("point") for the dot and annotate("text") for the label, plus four annotate("segment") calls to draw the border box. Coordinates must be positioned outside the plot area using coord_cartesian(clip="off"). 2) Stacked Bar Chart: Reverse stacking The B to Plasma differentiation chart needed Plasma cells on top visually (mature cells above precursors). In ggplot2: position_stack(reverse=TRUE) reverses the stacking order, but you also need guides(fill=guide_legend(reverse=TRUE)) to match the legend order. It helps to avoid visual confusion. 3) Network Graph: Proportional arrow size To make network arrows proportional to interaction strength edge weights determination is required: E(g)$arrow.size <- E(g)$weight * scaling_factor. The scaling factor (I used 2) determines visual range. If the factor is too small, differences disappear; if it is too large, arrows become cartoonish. When edges overlap (curved=0 for straight lines), bidirectional arrows look like one double-headed arrow, even though they are two separate directed edges with different weights. 4) Final Assembly When assembling, library(patchwork), library(cowplot) and library(magick) were used. All visualizations were created as part of the HackBio 2025 Data Visualization in Bio bioinformatics internship. #Bioinformatics #DataVisualization #HackBio #R #ggplot2 #Immunology #SingleCellRNAseq #InfectiousDisease #SystemsBiology
-
+4
-
In epidemiological research, Relative Risk (RR), Odds Ratio (OR), and Attributable Risk (AR) are commonly reported — but each answers a different question. Let’s look at a simple example: In a study on smoking and lung disease: • Among 100 smokers, 20 develop lung disease • Among 100 non-smokers, 10 develop lung disease 🔹 Relative Risk (RR) Risk in smokers = 20/100 = 0.20 Risk in non-smokers = 10/100 = 0.10 RR = 2 ➡️ Smokers are 2 times more likely to develop lung disease compared with non-smokers. 🔹 Odds Ratio (OR) OR ≈ 2.25 ➡️ Smokers have about 2.25 times higher likelihood of developing lung disease compared with non-smokers. 🔹 Attributable Risk (AR) AR = 0.20 − 0.10 = 0.10 (10%) ➡️ 10% of the disease risk is attributable to smoking. 📊 In simple terms: • RR → compares probabilities • OR → compares likelihood between groups • AR → quantifies the excess risk due to exposure Understanding these measures is essential for accurate interpretation of epidemiological and clinical research findings. #Biostatistics #Epidemiology #ClinicalResearch #EvidenceBasedMedicine #PublicHealth
-
Very often clinical trials will evaluate the averages of more than two treatments. One of the most fundamental tools statisticians use to evaluate if there is a difference between three or more treatments is the Analysis of Variance or ANOVA. The ANOVA is used to state if a statistical difference exists between the treatments under study - IT IS NOT USED TO IDENTIFY WHICH GROUPS ARE DIFFERENT FROM EACH OTHER. Stated another way - the ANOVA tells you that at least one group is different, but it does not specify which groups are different. The ANOVA does this by looking at the amount of variation between the groups and comparing it with the amount of variation within the groups. Kind of crazy that to evaluate averages (means) we look at the variance of the data right? This is a great example of why I always say that statistics is all about variance 😃. The real question the ANOVA test is trying to answer is: "Is the difference among groups greater than what is expected to be caused by chance?" If you are working to become a statistician or a statistical programmer OR if you are working to design a clinical study it is CRITICAL to learn about ANOVA. ANOVAs can: - Make clinical trials more cost-effective (you don't have to run multiple trials to compare all these various treatments against each other one at a time) - Help with dose selection (especially in phase 2 proof of concept studies when a company is still deciding between two doses) - Minimize bias (it is easier to show differences among multiple groups are due to one item (treatment) when compared together) If you would like to learn more about ANOVA and other statistical tests visit my website (link in bio) or message me directly. Happy Learning and Happy Wednesday
-
Relative risk looks simple until people start interpreting it incorrectly. I see this often with students reading papers, writing theses, or preparing for public health exams. RR answers one clean question: → How does the risk of an outcome compare between two groups? If RR = 1: → The risk is the same in both groups If RR > 1: → The exposed group has higher risk If RR < 1: → The exposed group has lower risk The common mistakes: → Confusing risk with odds ↳ RR compares risk, not odds. → Saying RR = 5 means “5% higher risk” ↳ It means 5 times the risk. → Ignoring the confidence interval ↳ If the 95% CI crosses 1, be careful with the interpretation. → Using RR when the study design does not allow direct risk estimation ↳ In many case-control studies, odds ratio is usually the correct measure. A simple memory cue: RR is best when you can calculate risk in both groups. That is why it fits naturally with cohort studies. Before interpreting RR, ask: → What is the exposed group? → What is the unexposed group? → What outcome was measured? → Does the confidence interval include 1? Small errors in interpreting RR can change how people read an entire paper. Which measure should I simplify next: odds ratio, hazard ratio, risk difference, or attributable risk? #Epidemiology #PublicHealth #ResearchMethods #Biostatistics
-
Single-cell data can now reveal which enhancers control which genes in specific cell types! Paper of the day, from Nature Portfolio! Many disease-associated variants lie outside protein-coding regions. The hard part is identifying which gene each regulatory variant affects, and in which cell type. The team developed scE2G, a family of supervised models that combines chromatin accessibility, genomic distance, promoter properties, and, for paired multiome data, RNA-accessibility correlations. It was trained on 10,342 enhancer-gene pairs tested by CRISPR. Against ten published single-cell methods, scE2G performed best across independent CRISPR, eQTL, and GWAS benchmarks. Its multiome version reached 14.9-fold eQTL enrichment, while cell-type-resolved bone marrow analysis found 76% more unique enhancer-gene links than pooled analysis. Applied across 45 cell types, the maps nominated candidate long-range links from the blood-trait variant rs7696969 to INPP4B and IL15. Great work by co-first authors Maya Sheth and Wei-Lin Qiu, with Jesse Engreitz and Robin Andersson, across the Broad Institute of MIT and Harvard, Stanford University, and the University of Copenhagen (Københavns Universitet)! Paper: https://lnkd.in/gpRftMSj The honest caveat: training included only 466 positive links, all from K562 cells, and performance in other cell types remained modest. Broader perturbation data across cell types will be essential. This connects directly with my current research: using AI to turn genomic data into testable biological hypotheses, not just more descriptive maps. Which disease area would benefit most from these maps: immune disease, neurodegeneration, cancer, or metabolic disease? #AIforScience #AI #MachineLearning #Science #Research #Genomics #SingleCellGenomics #MultiOmics #Epigenomics #GeneRegulation #Enhancers #CRISPR #FunctionalGenomics #HumanGenetics #NoncodingVariants #ComputationalBiology #Bioinformatics #PrecisionMedicine #AIforBiology #GWAS
-
Statistical Methods in Clinical Trials: 1)Analysis of Variance (ANOVA): Purpose: Compares means between multiple groups to determine significant differences. Application: Assessing treatment efficacy across different dosage levels or arms. SAS Syntax: PROC ANOVA with CLASS and MODEL statements. **SAS Syntax:** PROC ANOVA data=trial_data; CLASS treatment_group; MODEL efficacy_outcome = treatment_group; RUN; 2)T-tests: Purpose: Compares means between two independent groups. Application: Comparing treatment and control groups. SAS Syntax: PROC TTEST with CLASS and VAR statements. **SAS Syntax:** PROC TTEST data=trial_data; CLASS treatment_group; VAR efficacy_outcome; RUN; 3)Chi-square test: Purpose: Analyzes categorical data to find significant associations. Application: Comparing response rates between treatment groups. SAS Syntax: PROC FREQ with TABLES statement for CHISQ test. **SAS Syntax:** PROC FREQ data=trial_data; TABLES treatment_group * response_category / CHISQ; RUN; 4)Survival Analysis: Purpose: Analyzes time-to-event data, such as survival rates. Application: Assessing progression-free survival or overall survival. SAS Syntax: PROC LIFETEST for Kaplan-Meier curves and PROC PHREG for Cox proportional hazards model. **Kaplan-Meier Curve SAS Syntax:** PROC LIFETEST data=trial_data; TIME time_to_event*event(0); STRATA treatment_group; RUN; **Cox Proportional Hazards Model SAS Syntax:** PROC PHREG data=trial_data; CLASS treatment_group; MODEL time_to_event*event(0) = treatment_group; RUN; 5)Linear Regression: Purpose: Evaluates relationships between predictor variables and continuous outcomes. Application: Assessing impact of covariates on efficacy outcomes. SAS Syntax: PROC REG with MODEL statement. **SAS Syntax:** PROC REG data=trial_data; MODEL efficacy_outcome = predictor_variable1 predictor_variable2; RUN; 6)Logistic Regression: Purpose: Analyzes binary outcome variables (e.g., response vs. non-response). Application: Analyzing categorical outcomes in clinical trials. SAS Syntax: PROC LOGISTIC with CLASS and MODEL statements for binary outcomes. **SAS Syntax:** PROC LOGISTIC data=trial_data; CLASS treatment_group; MODEL binary_outcome(event='1') = treatment_group; RUN; 7)Repeated Measures Analysis: Purpose: Examines data collected over time, considering within-subject changes. Application: Assessing longitudinal efficacy outcomes across multiple time points. SAS Syntax: PROC MIXED with MODEL statement and RANDOM statement for repeated measures. **SAS Syntax:** PROC MIXED data=trial_data; CLASS subject_id timepoint; MODEL efficacy_outcome = treatment_group timepoint treatment_group*timepoint / SOLUTION; RANDOM subject_id; RUN; #sas #clinicalsas #clinicalsasprogrammimg #biostatastic #statasticalprogrammimg #cdisc #adam #tlf #sdtm #clinicaltrial #r #sasprogrammimg