/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Multimodal

1296 articles tagged with Multimodal

Latest Trending
Mastodon discussion 5d ago

University of Texas study, for blind or low vision participants: Living in an AI-filtered reality: Participant recruitme...

University of Texas study, for blind or low vision participants: Living in an AI-filtered reality: Participant recruitment & eligibility screening form docs.google.com/forms/d/e/1F...

Google Multimodal
9
GitHub Trending repo 5d ago

MoonshotAI/PerceptionBench: PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Multimodal
60
Mastodon discussion 5d ago

Ant Group's LingBot subsidiary has released six open-source embodied AI models at WAIC 2026, covering vision, video, spa...

Ant Group's LingBot subsidiary has released six open-source embodied AI models at WAIC 2026, covering vision, video, spatial perception, manipulation, world models and world action...

Multimodal Open Source Robotics
9
Mastodon discussion 6d ago

Medtronic introduces Touch Surgery Aide, an AI compute platform for the OR that uses real-time multimodal AI and compute...

Medtronic introduces Touch Surgery Aide, an AI compute platform for the OR that uses real-time multimodal AI and computer vision to support surgical workflows and decision-making. ...

Multimodal
9
AlternativeTo tool 6d ago

Moonshot AI launches Kimi K3, a 2.8T parameter open MoE model with vision support

Last week, Beijing based AI startup Moonshot AI launched Kimi K3, a 2.8 trillion parameter Mixture of Experts model described as the first open model in its class and the largest o...

Anthropic Multimodal
9
NewsData.io news 6d ago

ANNA Money Accelerates Vision to Build a Comprehensive AI Platform for Small Businesses with Business Data Group and UK Business Forums Acquisition

ANNA Money , the AI-powered business account that does your taxes, today announces the acquisitions of Business Data Group (BDG) and UK Business Forums (UKBF), marking another majo...

Multimodal
21
GNews news Jul 22

Sony and Mitsubishi Set Up AI Vision Joint Venture for Factories

Mitsubishi Electric Corp. and Sony Group Corp. said they will set up a joint venture to provide factory automation services by combining their expertise in artificial intelligence ...

Multimodal
18
NewsData.io news Jul 22

Mitsubishi Electric, Sony to form AI vision sensor venture for factory automation

Mitsubishi Electric, Sony to form AI vision sensor venture for factory automation

Multimodal
21
Papers with Code paper Jul 22

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful...

Multimodal
21
Papers with Code paper Jul 22

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulat...

Multimodal
21
Mastodon discussion Jul 21

Ethical, Low-Power, and On-Device: Studio Atelico’s Vision for the Future of Generative AIhttps://www.gamesindustry.biz/...

Ethical, Low-Power, and On-Device: Studio Atelico’s Vision for the Future of Generative AIhttps://www.gamesindustry.biz/ethical-low-power-and-on-device-studio-atelicos-vision-for-t...

Multimodal
18
Dev.to tutorial Jul 21

Kimi K3 API Guide: Reasoning, Tool Calling, Structured Output, and Vision

If you are testing Kimi K3 through an OpenAI-compatible API, there are a few details worth knowing...

OpenAI Multimodal API
12
Mastodon discussion Jul 21

An issue about #AI is that we probably have no clear vision of what we want from it. We don't try to solve problems and ...

An issue about #AI is that we probably have no clear vision of what we want from it. We don't try to solve problems and just applying AI everywhere hoping that it will stick.It may...

Multimodal
9
Mastodon discussion Jul 21

@jeffjarvis 3/The Currency:LLM: Discrete text/multimodal tokens. No body, no hunger, no pain—just pure mathematical rela...

@jeffjarvis 3/The Currency:LLM: Discrete text/multimodal tokens. No body, no hunger, no pain—just pure mathematical relationships between symbols.Human Brain: Action potentials, ne...

LLM Multimodal
27
Papers with Code paper Jul 21

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements...

Multimodal
21
Papers with Code paper Jul 21

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communic...

Multimodal
21
NewsData.io news Jul 20

NodeMeta Expands Its Multi-Utility Web3 Vision With OneKey, TrustScan AI and NTE

NodeMeta is a community-driven Web3 infrastructure and utility ecosystem powered by NTE. It is developing a connected platform involving nodes, staking, AI-enab

Multimodal
21
Hugging Face model Jul 20

baseten/GLM-5.2-Vision-NVFP4 - image-text-to-text

Hugging Face model: baseten/GLM-5.2-Vision-NVFP4

Multimodal
51
GitHub Trending repo Jul 20

MalaikaTariq24/multimodal-image-studio: 🎨 Text-to-Image AI App | DecodeLabs Project 3

🎨 Text-to-Image AI App | DecodeLabs Project 3

Image Generation Multimodal
32
GNews news Jul 20

Hesai Showcases Kosmo Spatial Intelligence Platform and Robotic Lidar at WAIC 2026, Underscoring Its Physical AI Vision

Hesai Technology (NASDAQ: HSAI; HKEX: 2525) recently participated in the World Artificial Intelligence Conference (WAIC) 2026 in Shanghai, presenting Kosmo,

Multimodal
18
Mastodon discussion Jul 20

Inkling releases open-weights 975B multimodal model designed for fine-tuning#AI #AutomationSource: Product Hunt AIhttps:...

Inkling releases open-weights 975B multimodal model designed for fine-tuning#AI #AutomationSource: Product Hunt AIhttps://www.producthunt.com/products/tinker-2

Multimodal
18
NewsData.io news Jul 20

Zoho's Chief Scientist on Bharat as Vishwaguru: 'The vision is to revive an era of...'

Sridhar Vembu says the post-World War II model of technology transfer helped nations become self-reliant, unlike today's intellectual property-driven globalisation

Multimodal
21
Papers with Code paper Jul 20

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most app...

Multimodal
21
Papers with Code paper Jul 20

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires applicat...

Multimodal
21
« Previous Page 3 of 54 (1296 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available