Open Source Innovation Platforms

Explore top LinkedIn content from expert professionals.

  • View profile for Bhavishya Pandit

    Turning AI into enterprise value | $20 M in Business Impact | Speaker - MHA/IITs/IIMs/NITs | Google AI Expert | 50 Million+ views | MS in ML - UoA

    85,997 followers

    Meta went bonkers with this new open-source ASR that works for 1,600+ languages! 🤯 Now, businesses can reach customers in their native tongue, even in low-resource regions, without building ASR from scratch. → Fully open-source, supporting 500+ languages never covered by any ASR before → Trained on 4.3M hours of multilingual speech (1,600+ languages) → Best part: Works zero-shot on languages never seen during training How? Two breakthroughs: Dual-decoder architecture:  • CTC decoder for low-latency, real-time use  • LLM-ASR decoder (Transformer-based) for high-accuracy, context-aware transcription In-context learning: Just 5–10 speech-text examples at inference time, let it transcribe any new language even if the model was never trained on it. Even more surprising: → On FLEURS-81, Omnilingual ASR beats Whisper on 65/81 languages—including 24 of the world’s top 34 most spoken languages → Robust to noise: CER stays <10 even in the noisiest 5% of field recordings → Scales from edge to cloud: 300M (mobile) → 7B (max accuracy) But the real shift isn’t scale, it’s agency. Communities can now extend ASR to their own language with minimal data, compute, or expertise. Check out the carousel to know how it works in simple terms and what the challenges are in detail. Question for you: When building voice tech for underserved languages, do you prioritise zero-shot generalisation or lightweight fine-tuning and why? Follow me, Bhavishya Pandit, for honest takes on AI tools that actually work 🔥 P.S. Model card, inference code, and datasets in the first comment.

  • View profile for Leslie Teo

    Building open, local AI for Southeast Asia | Senior Director, AI Products at AI Singapore (SEA-LION) | Member, UN Scientific Panel on AI | Former GIC Chief Economist

    9,003 followers

    🌟 𝗧𝗵𝗲 𝗕𝗲𝘀𝘁 𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴 𝗠𝗼𝗱𝗲𝗹𝘀 𝗳𝗼𝗿 𝗦𝗼𝘂𝘁𝗵𝗲𝗮𝘀𝘁 𝗔𝘀𝗶𝗮 (𝗮𝗻𝗱 𝗢𝗽𝗲𝗻 𝗧𝗼𝗼!) 🌟 Most embedding models treat Southeast Asian languages as a byproduct. The result: search and retrieval systems with blind spots — whole chunks of information in Thai, Malay, or Indonesian that just don't surface when they should. Get embeddings wrong and every system built on top — RAG, memory, agents — inherits those blind spots. We built the fix. Three open embedding models designed specifically for SEA. Top of SEA-BED (https://lnkd.in/gh7zEHdu) — the benchmark built for Southeast Asian language understanding — beating multilingual-E5-large, BGE-M3, and Qwen's 8B model at just 600M parameters. 13x smaller, better results. Beats closed models too (see leaderboard for details.) Three models, two checkpoints:  🔹 𝗦𝗘𝗔-𝗟𝗜𝗢𝗡-𝗘𝟱-𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴-𝟲𝟬𝟬𝗠 — fine-tuned on multilingual-E5 · 512 token context  🔹 𝗦𝗘𝗔-𝗟𝗜𝗢𝗡-𝗠𝗼𝗱𝗲𝗿𝗻𝗕𝗘𝗥𝗧-𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴-𝟯𝟬𝟬𝗠 — ModernBERT + Gemma 3 tokenizer · 8k context  🔹 𝗦𝗘𝗔-𝗟𝗜𝗢𝗡-𝗠𝗼𝗱𝗲𝗿𝗻𝗕𝗘𝗥𝗧-𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴-𝟲𝟬𝟬𝗠 — ModernBERT + Gemma 3 tokenizer · 8k context Both 300M and 600M checkpoints released for developers to fine-tune for their own use cases. We chose ModernBERT for its multilingual architecture and paired it with the Gemma 3 tokenizer — which handles SEA scripts with far less token fragmentation than most alternatives. Trained on 1–2T tokens from the region, covering Thai, Malay, Indonesian, Tagalog, Vietnamese and other major Southeast Asian languages. Glad to be able to include Tetum, ASEAN's newest member. The SEA-LION-E5-Embedding-600M punches hardest across the board. But the point isn't any single model — it's that SEA is the default, not the afterthought. While keeping strong English and Chinese performance too (both very much SEA languages 🙂). 𝗠𝗼𝗱𝗲𝗹𝘀: https://lnkd.in/gPhSdXRN 𝗕𝗹𝗼𝗴: https://lnkd.in/gyHFpDPx Special thanks to Raymond_ Ng and Peerat Limkonchotiwat for leading this work, Zhi_Hao Ong for building the embedding demo app, and Wuttikorn Ponwitayarat from Vidyasirimedhi Institute of Science and Technology (VISTEC) for developing the SEA-BED benchmark (https://lnkd.in/gengHyDt). Built with the support of Singapore NRF, IMDA, and MDDI. Thanks to our partners at Amazon Web Services (AWS), Google, and NVIDIA for deep engineering support. William Tjhi Mark Pereira David Ong Esther Choa Darius Liu, CFA, CAIA Jian Gang Ngui Abigail Toh Adwin Chan Rachel Liew Michelle Loh Pratyusha Mukherjee Richard Goh Dr. Deb Goswami Christopher Low Cathy C. Sarana Nutanong Shaowei Ying Ofir Shalev Tim Rosenfield #AI #AISG #AISingapore #MachineLearning #NLP #OpenSource  #SoutheastAsia #Gemma #Qwen

  • View profile for Arpit Singh

    GTM, AI & Outbound | LinkedIn Content & Social Selling for high-growth agencies, AI/SaaS startups & consulting businesses | Open for collaborations

    36,953 followers

    75% of internet users don't speak English. Yet most B2B sites are English-only. And then we wonder why international expansion feels complicated. It’s not complicated. It’s just operationally painful. For years, going global meant: Translator. Developer. SEO specialist. Weeks of coordination. So companies postponed it. But if you’re on WordPress, Shopify, Webflow, Squarespace or almost any CMS… You can make your site multilingual in minutes. That’s what I found interesting about Weglot. Install it once, and your site is instantly translated with AI. Behind the scenes, it also handles: • Language-specific URLs • Hreflang tags • Translated metadata • Automatic updates when content changes In other words, multilingual SEO without turning it into a dev project. You can refine key pages manually, control tone, and manage everything from one dashboard. Which changes the decision entirely. Instead of asking: “Are we ready to expand?” You ask: → “Which market should we test next?” If your website only speaks one language, you are not limiting ambition. You are limiting access. And access is leverage. Want to see how this works in practice? Try it here: https://lnkd.in/enXEbGFS If removing the technical friction made expansion easy… Which country would you test first?

  • View profile for Pan Wu
    Pan Wu Pan Wu is an Influencer

    Senior Data Science Manager at Meta

    52,270 followers

    Generative AI is transforming how people learn and work—but only if it speaks your language. Most AI features are still built English-first, and “just translate it” rarely delivers a great experience. Idioms, domain-specific terms, and cultural context often get lost when translation is treated as an afterthought. In a recent engineering blog, Udemy shares how they approached this challenge and built a framework for localizing generative AI features from the ground up. The team outlines three strategies along a spectrum of complexity. At one end is a Translation Management System (TMS): translate user input to English, run it through the LLM, then translate the output back. It’s fast to ship and offers broad coverage, but comes with tradeoffs in latency and nuance. At the other end is a Multilingual LLM System (MLS), where the model processes and generates directly in each language using multilingual prompting, cross-lingual embeddings, and optional fine-tuning. This delivers higher quality, but is more complex to build. In between sits a hybrid approach—routing simpler queries through TMS and high-stakes interactions through multilingual models—allowing teams to move fast while investing deeply where it matters most. What stands out is how they treat localization as a platform problem: core interfaces are designed to switch between TMS and MLS without major rewrites. Safety and compliance are validated per language, rather than assumed to generalize from English. And every new language follows a repeatable playbook: start with TMS, learn from real usage, then decide whether it’s worth upgrading to a fully multilingual system. The results are impressive: the team was able to go from concept to production for the Japanese market in under three months, and adding a new language now takes less than 25% of the original effort. The takeaway? Localization isn’t a tax on your AI roadmap—it’s a multiplier. Start broad, go deeper where the data justifies it, and design your system so scaling globally becomes the default. #DataScience #DecisionMaking #LLM #Translation #Platform #SnacksWeeklyonDataScience – – –  Check out the "Snacks Weekly on Data Science" podcast and subscribe, where I explain in more detail the concepts discussed in this and future posts:    -- Spotify: https://lnkd.in/gKgaMvbh   -- Apple Podcast: https://lnkd.in/gFYvfB8V    -- Youtube: https://lnkd.in/gcwPeBmR https://lnkd.in/ggZxBDYj

  • View profile for Alex Issakova

    Helping B2B commercial teams turn GenAI confusion into practical use cases, workflows and confident everyday adoption | Ex-Senior Exec, Silicon Valley (IPO’d) | In AI since 2013 | Keynote Speaker | Author of The Roadmap

    34,167 followers

    Switzerland’s Ethical GPT Alternative Is Finally Live 🚀 Built in the Alps. Powered by green energy. Trained in 1,000+ languages. And open to everyone. Apertus, Switzerland’s first open-source GPT-4-scale LLM has officially launched. And it’s rewriting the rules of AI development. While OpenAI, Anthropic, and Google dominate headlines with closed models, Switzerland just quietly released something different: a multilingual, energy-efficient, fully transparent alternative. And they’re giving it away. 🔐 True Digital Independence Not another hype model. A statement: Built by public institutions (ETH Zurich, EPFL) Apache 2.0 licensed — completely open Full access to weights, training data, and recipes No Big Tech lock-in. No black boxes. No strings attached This is AI infrastructure that everyone can inspect, improve, and trust. 🌱 AI Powered by the Alps Apertus was trained using Switzerland’s carbon-neutral supercomputers — running on 100% renewable alpine hydropower. Sustainable AI is not a buzzword here Its energy footprint is radically smaller than GPT-4 and Claude Proving you can build world-class models without draining global energy grids 🌍 AI That Speaks 1,000+ Languages Unlike most commercial models, Apertus was trained, not just fine-tuned, on over 1,000 languages. That means underserved communities finally get native-quality AI: Regional dialects Indigenous languages True multilingual parity AI shouldn’t just work for English speakers. Now, it doesn’t have to. ⚖️ Built With Ethics First, Not After Apertus is EU AI Act-ready from day one: Training data is fully documented Respects copyright and website opt-out signals Privacy isn’t bolted on — it’s baked in Transparency isn’t optional here. It’s the entire point. But Here’s the Catch… Apertus isn’t perfect — yet. It lags GPT-4 and Claude on advanced reasoning, coding, and multi-step problem solving It’s missing some ecosystem perks — no plugins, memory, or multimodal features Running the 70B model requires serious compute; the smaller 8B model is more accessible but less capable Still, for an open model just launching? It’s an impressive foundation — and it’s only going to improve. 💡 Why This Matters Apertus proves something critical: LLMs don’t have to be closed, extractive, or environmentally destructive. This is a blueprint for the future: Public ownership Academic freedom Energy responsibility Global accessibility Open. Transparent. Ethical. Exactly what AI should be. ♻️ If this resonated, share it with your network. 🔔 Follow Alex Issakova for more reflections on AI and leadership. 

  • View profile for Tom Aarsen

    🤗 Sentence Transformers & NLTK maintainer, MLE @ Hugging Face

    20,872 followers

    ModernBERT goes MULTILINGUAL! One of the most requested models I've seen, The Johns Hopkins University's CLSP has trained state-of-the-art massively multilingual encoders using the ModernBERT architecture: mmBERT. Model details: - 2 model sizes: 42M non-embed (140M total) and 110M non-embed (307M total) - Uses the ModernBERT architecture, but with the Gemma2 multilingual tokenizer (so: flash attention, alternating global/local attention, unpadding/sequence packing, etc.) - Maximum sequence length of 8192 tokens, on the high end for encoders - Trained on 1833 languages using DCLM, FineWeb2, and many more sources - 3 training phases: 2.3T tokens pretraining on 60 languages, 600B tokens mid-training on 110 languages, and 100B tokens decay training on all 1833 languages. - Also uses model merging and clever transitions between the three training phases. - Both models are MIT Licensed, and the full datasets and intermediary checkpoints are also publicly released Evaluation details: - Very competitive with ModernBERT at equivalent sizes on English (GLUE, MTEB v2 English after finetuning) - Consistently outperforms equivalently sized models on all Multilingual tasks (XTREME, classification, MTEB v2 Multilingual after finetuning) - In short: beats commonly used multilingual base models like mDistilBERT, XLM-R (multilingual RoBERTa), multilingual MiniLM, etc. - Additionally: the ModernBERT-based mmBERT is much faster than the alternatives due to its architectural benefits. Easily up to 2x throughput in common scenarios. Check out the full blogpost with more details. It's super dense & gets straight to the point: https://lnkd.in/ebqTK3JS Based on these results, mmBERT should be the new go-to multilingual encoder base models at 300M and below. Do note that the mmBERT models are "base" models, i.e. they're currently only trained to perform Mask Filling. They'll need to be finetuned for downstream tasks like semantic search, classification, clustering, etc. I'm very much looking forward to seeing embedding models based on mmBERT! Great work by Marc Marone, Orion Weller, and the rest of the team at JHU!

  • View profile for Chandra Sekhar

    I simplify AI for everyone | 50K+ Followers | Top 1% Linkedin India | Senior AI Engineer | Agentic AI Trainer | Full Stack Gen AI Trainer | Corporate Trainer

    50,727 followers

    🌍 RAG System Design Question (the kind that filters senior engineers) "A company wants a multilingual RAG system supporting 20 languages across millions of documents. Accuracy and retrieval speed are critical. How would you design the end-to-end architecture?" Most people jump straight to "use a vector DB." But the real signal is how you handle scale + multilingual + latency together. 👇 Here's the end-to-end architecture I'd propose: 1️⃣ Ingestion & Language Handling → Detect language per document (fastText / CLD3) → Clean, chunk semantically (300–500 tokens with overlap), preserve metadata: language, source, timestamp → Store raw + normalized text for auditing 2️⃣ Embeddings — go multilingual, not per-language → Use a multilingual embedding model (e.g. multilingual-e5, BGE-M3, Cohere multilingual) so all 20 languages share ONE semantic space → This lets a query in Hindi retrieve a relevant doc in English — true cross-lingual retrieval → Avoid 20 separate indexes; it kills recall and ops sanity 3️⃣ Vector Store at Scale → Millions of docs → use ANN indexes (HNSW / IVF-PQ) in Qdrant, Milvus, or pgvector → Shard + filter by metadata (language, tenant, recency) to shrink the search space → faster retrieval → Add a hybrid layer: BM25 (keyword) + dense vectors → fuse with RRF for accuracy on names, codes, acronyms 4️⃣ Retrieval Pipeline (accuracy lever) → Stage 1: fast ANN recall top-100 → Stage 2: cross-encoder re-ranker (multilingual) → top-5 → Optional query translation/expansion for low-resource languages 5️⃣ Generation → Pass re-ranked context to a multilingual LLM → Answer in the user's query language, cite sources → Add grounding/guardrails to reduce hallucination 6️⃣ Speed & Reliability → Cache embeddings + frequent queries (semantic cache) → Async batch ingestion, precompute embeddings → Monitor: retrieval recall@k, latency p95, per-language accuracy ======================================== I’ve covered questions like these in my AI Engineering Interview Master Bundle, a comprehensive set of 22 courses designed for real interview prep. Explore the full guide here → https://lnkd.in/gqFkWZd4

  • View profile for Arunkumar Palanisamy

    Integration Architect → Senior Data Engineer | AI/ML | 19+ Years | AWS, Snowflake, Spark, Kafka, Python, SQL | Retail & E-Commerce

    4,512 followers

    𝐀𝐫𝐞 𝐘𝐨𝐮 𝐒𝐭𝐢𝐥𝐥 𝐏𝐚𝐲𝐢𝐧𝐠 𝐟𝐨𝐫 𝐚 𝐅𝐫𝐨𝐧𝐭𝐢𝐞𝐫 𝐀𝐏𝐈 𝐖𝐡𝐞𝐧 𝐚𝐧 𝐎𝐩𝐞𝐧-𝐒𝐨𝐮𝐫𝐜𝐞 𝐌𝐨𝐝𝐞𝐥 𝐖𝐨𝐮𝐥𝐝 𝐃𝐨 𝐭𝐡𝐞 𝐉𝐨𝐛? You don't need a proprietary API to build serious AI anymore. Nine open models cover almost every use case: chatbots, coding, multilingual, on-device: most with commercial licenses. The hard part isn't access. It's picking the right one. 𝐆𝐞𝐧𝐞𝐫𝐚𝐥 𝐩𝐮𝐫𝐩𝐨𝐬𝐞: 1. Llama 3 (Meta): 8B and 70B. Strong general performance, actively maintained, commercial use allowed. The de facto default of open models. When you're not sure where to start, Llama's ecosystem and community make it the safe first pick. 2. Falcon 7B/40B (TII): Trained on high-quality data, strong instruction following, Apache 2.0. Proved a non-US lab could ship a top-tier open model. Permissive license makes it enterprise-friendly from day one. Efficiency and cost: 3. Mistral 7B / Mixtral 8x7B: High performance, long context (32K), Apache 2.0, great on-device and cloud. Mixtral's MoE design delivers big-model quality while activating only a fraction of parameters. Performance without the full compute bill. 4. Stable LM 2 (Stability AI): 1.6B and 12B. Instruction tuned, permissive license, fast inference. When speed and local deployment matter more than maximum capability, a small fast model is the pragmatic choice. Small but mighty: 5. Phi-3 (Microsoft): Mini, Small, Medium. Excellent reasoning and coding, MIT licensed, runs on CPUs. Proves curated training data beats raw size. Punches far above its parameter count. Best for coding assistants, educational tools, and resource-constrained setups. 6. Gemma (Google DeepMind): 2B and 7B. Built from Gemini research, strong safety and alignment, optimized for inference. When responsible behavior matters and resources are tight. Google's safety work in a small open package. Multilingual: 7. BLOOM (BigScience): 560M to 176B. 46+ languages, radically transparent, community-driven. If your users aren't English-first, this is a serious option. Breadth of languages is the edge no other open model matches. 8. Qwen (Alibaba Cloud): 1.8B to 72B. Strong reasoning, multilingual, long context (32K+), Apache 2.0. Quietly one of the strongest open families, especially for multilingual and reasoning tasks. Often underrated in Western stacks. Code-specialized: 9. Code Llama (Meta): 7B, 13B, 34B. Fine-tuned specifically for code, multiple languages, long context (16K), free for commercial use. A coding specialist beats a generalist at coding tasks every time. If you're building dev tools, the switch is worth it. No single best open model. Only the right one for your constraint size, language, license, task, or compute budget. Match the model to the job, not to the leaderboard. Which one is running in your stack and did you pick it for licensing, performance, or size? ♻️ Repost this to help your network get started ➕ Follow Arunkumar for more #OpenSource #LLM #AIEngineering

  • View profile for Asif Razzaq

    Founder @ Marktechpost (AI Dev News Platform) | 1 Million+ Monthly Readers

    38,177 followers

    CMU Researchers Release Pangea-7B: A Fully Open Multimodal Large Language Models MLLMs for 39 Languages A team of researchers from Carnegie Mellon University introduced PANGEA, a multilingual multimodal LLM designed to bridge linguistic and cultural gaps in visual understanding tasks. PANGEA is trained on a newly curated dataset, PANGEAINS, which contains 6 million instruction samples across 39 languages. The dataset is specifically crafted to improve cross-cultural coverage by combining high-quality English instructions, machine-translated instructions, and culturally relevant multimodal tasks. In addition, to evaluate PANGEA’s capabilities, the researchers introduced PANGEABENCH, an evaluation suite spanning 14 datasets covering 47 languages. This comprehensive evaluation provides insight into the model’s performance on both multimodal and multilingual tasks, showing that PANGEA outperforms many existing models in multilingual scenarios. PANGEA was developed using PANGEAINS, a rich and diverse dataset that includes instructions for general visual understanding, document and chart question answering image captioning, and more. The dataset was designed to address the major challenges of multilingual multimodal learning: data scarcity, cultural nuances, catastrophic forgetting, and evaluation complexity. To build PANGEAINS, the researchers employed several strategies: translating high-quality English instructions, generating culturally aware tasks, and incorporating existing open-source multimodal datasets. The researchers also developed a sophisticated pipeline to filter culturally diverse images and generate detailed multilingual and cross-cultural captions, ensuring that the model understands and responds appropriately in different linguistic and cultural contexts... Read the full article here: https://lnkd.in/gV5Bnac5 Paper: https://lnkd.in/gcxNAQVy Model on Hugging Face: https://lnkd.in/gMKtJ83z Project Page: https://lnkd.in/gP7vWBkb Listen to the podcast on Pangea-7B created with the help of NotebookLM and, of course, with the help of our team, who generated the prompts and entered the right information: https://lnkd.in/gTBvt9NG Machine Learning Department at CMU Xiang Yue Yueqi Song Seungone Kim Jean de Dieu Nyandwi Simran Khanuja Lintang Sutawika Sathyanarayanan Ramamoorthy Graham Neubig

  • View profile for Olga R.

    Medical Strategist | Health Geek | Ghostwriter | Bridging Yesterday’s Lessons to Tomorrow’s Medicine

    3,592 followers

    The Swiss have already beaten us in longevity. Now they’re aiming to outpace us in AI—openly. ETH Zurich, EPFL, and CSCS just announced something radical: They’re releasing a fully open-source large language model (LLM)—code, weights, and training data—built on carbon-neutral public infrastructure. Multilingual. Regulation-aligned. And powerful enough to rival major commercial models (up to 70B parameters). It’s not built for clicks. It’s not locked in a walled garden. It’s built for the public good. Kind of like when Volvo invented the 3-point seatbelt in 1959—and instead of patenting it, they gave it away. Because some tools are too important to keep private. Here’s why this matters to anyone in healthcare, research, or digital health: 1. Democratized AI access Most advanced LLMs are locked behind paywalls or corporate APIs. This one isn’t. ✅ Clinicians, researchers, and startups can build tools faster and cheaper ✅ Hospitals can run models locally—no need to send data to Big Tech ✅ Small players now have enterprise-level AI capabilities 2. Trustworthy AI Built under Swiss + EU regulations, this model is transparent, ethical, and compliant from the start—something U.S. tools are still wrestling with. This is AI you can trust in medicine, where accuracy and safety matter. 3. Multilingual = Global care Trained in 1,500+ languages. Fluent in over 1,000. That changes everything. 🗣️ Global health equity 🌍 Medical translation 💬 Patient-facing tools in diverse communities This is the infrastructure for inclusive care. 4. A broader movement The Swiss are known for longevity, health innovation, and public trust. Now they’re applying those same values to AI—ethical, open, and pro-human. It’s the opposite of the closed, extractive systems clinicians are burning out in. 5. A wake-up call If you work in health tech, pharma, or clinical research and you're not preparing for open, multilingual, regulator-aligned AI— You're already behind. As Dr. Dante Morra said this week: “The next disruption in healthcare isn’t clinical—it’s personal.” And if your tools aren’t ready to meet patients where they are? They’ll build new ones. The Swiss just handed us one.

Explore categories