Meta

2025Meta

Action100M v1.0

Open Source
Robotics
CV

Action100M is a large-scale video action dataset developed by Meta, designed to advance research in action recognition and understanding. It comprises 100 million YouTube video clips, each annotated with hierarchical action labels, enabling the study of complex human activities in diverse contexts.

100 million YouTube video clips
Hierarchical action annotations
Diverse human activities
PricingFree
Version1.0
2025Meta

Llama 3.1

Open Source
Open Source
LLM

Meta's latest open-source large language model with 405B parameters. Offers state-of-the-art performance on benchmarks while being freely available for commercial use.

405B, 70B, and 8B parameter versions available
128K token context length
Multilingual support for 8 languages
PricingFree and open source
RegionUnited States
2026Meta

Muse Spark 1.1 v1.1

Paid
Featured
Multimodal AI
Agentic AI

Muse Spark 1.1 is a multimodal reasoning model designed for agentic tasks, featuring enhanced capabilities in tool use, computer interaction, and complex coding. Developed by Meta Superintelligence Labs, it supports long-context management and efficient multi-agent orchestration for enterprise-grade workloads.

1-million-token context window
Autonomous multi-agent orchestration
Computer-use automation and direct interface interaction
Version1.1
RegionUnited States
2026Meta

FFmpeg at Meta: Media Processing at Scale

Open Source
Infrastructure
ML

Meta has integrated FFmpeg, an open-source multimedia framework, to enhance media processing capabilities across its platforms. By collaborating with FFmpeg developers, FFlabs, and VideoLAN, Meta has contributed to the development of features like threaded multi-lane transcoding and real-time quality metrics, which have been incorporated into the upstream FFmpeg project.

Threaded multi-lane transcoding
Real-time quality metrics
Efficient multi-output encoding
PricingFree
RegionUnited States
2025Meta

Perception Encoder Audiovisual (PE-AV) v1.0

Open Source
CV
ML

Meta's Perception Encoder Audiovisual (PE-AV) is a state-of-the-art multimodal model that embeds audio, video, audio-video, and text into a joint embedding space. Trained using contrastive learning on approximately 100 million audio-video pairs with text captions, PE-AV enables powerful cross-modal retrieval and understanding across audio, video, and text modalities. ([marktechpost.com](https://www.marktechpost.com/2025/12/22/meta-ai-open-sourced-perception-encoder-audiovisual-pe-av-the-audiovisual-encoder-powering-sam-audio-and-large-scale-multimodal-retrieval/?utm_source=openai))

Embeds audio, video, and text into a unified embedding space
Trained on 100 million audio-video pairs with text captions
Enables cross-modal retrieval and understanding
PricingFree
Version1.0
2026Meta

Muse Code vbeta

Free
Featured
AI Coding Agent
Software Engineering

Muse Code is a terminal-based coding agent powered by the Muse Spark 1.2 model, designed to handle complex software engineering tasks across large repositories. It features persistent background agents, repository-scale execution, and built-in verification to automate planning, coding, and debugging workflows.

Persistent async background agents for reduced latency
Restart-safe runtime with local event logging
Built-in skills: /plan, /grill, and /goal
Versionbeta
RegionUnited States

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode