Skip to main content

Model profile

LFM2.5-VL-1.6B-Extract

LiquidAICurrentReleased May 26, 2026
Data verified
Overall Score
Unranked
Arena Elo
Not listed
Eligible category ranks
0of 8
Price (1M tokens)
Not listedAPI pricing
Speed
Not listed
Context
128K

Evidence coverage

15 of 323 tracked benchmarks are published. 3 are verified and 12 provisional. 6 of 8 categories are measured.

Updated July 23, 2026Methodology
Published / tracked
15 / 323
Verified
3
Provisional
12
Categories with evidence
6 / 8

Evidence by category

  • Agentic1 benchmark
    Reported
  • Coding1 benchmark
    Reported
  • Reasoning2 benchmarks
    Reported
  • Knowledge6 benchmarks
    Reported
  • Math0 benchmarks
    Not measured
  • Multilingual0 benchmarks
    Not measured
  • Multimodal4 benchmarks
    Mixed evidence
  • Inst. Following1 benchmark
    Reported
Open WeightSelf-hostNon-Reasoning
Confidence:
Low
1-6b

BenchLM is tracking LFM2.5-VL-1.6B-Extract, but this profile is currently excluded from the public leaderboard because it still lacks enough non-generated benchmark coverage to rank safely. Only non-generated public benchmark rows appear below.

LFM2.5-VL-1.6B-Extract is a open weight model with a 128K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.

LFM2.5-VL-1.6B-Extract sits inside the LFM2.5-VL Extract family alongside LFM2.5-VL-450M-Extract. This profile currently has 15 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 77.4483.93

  1. Claude Mythos 5
    Anthropic
    #183.93
    Claude Mythos 5 is #1 with a score of 83.93.
    Compare
  2. Claude Fable 5
    Anthropic
    #283.68
    Claude Fable 5 is #2 with a score of 83.68.
    Compare
  3. GPT-5.6 Sol
    OpenAI
    #381.96
    GPT-5.6 Sol is #3 with a score of 81.96.
    Compare
  4. Kimi K3
    Moonshot AI
    #480.96
    Kimi K3 is #4 with a score of 80.96.
    Compare
  5. Claude Opus 4.8
    Anthropic
    #578.34
    Claude Opus 4.8 is #5 with a score of 78.34.
    Compare
  6. Muse Spark 1.1
    Meta
    #677.44
    Muse Spark 1.1 is #6 with a score of 77.44.
    Compare
  7. LFM2.5-VL-1.6B-ExtractCurrent model
    LiquidAI
    UnrankedNot measured
    LFM2.5-VL-1.6B-Extract is Unranked with a score of Not measured.

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

No eligible category percentile is available from the published evidence yet.

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticWeight 22%1 benchmarkReportedScore pending
CodingWeight 20%1 benchmarkReportedScore pending
ReasoningWeight 17%2 benchmarksReportedScore pending
KnowledgeWeight 12%6 benchmarksReportedScore pending
MathWeight 5%0 benchmarksNot measuredNot measured
MultilingualWeight 7%0 benchmarksNot measuredNot measured
MultimodalWeight 12%4 benchmarksMixed sourcesScore pending
Inst. FollowingWeight 5%1 benchmarkReportedScore pending

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic1 benchmark
τ²-bench resultsReported

τ²-Bench Tool-Agent-User Evaluation

8.5%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Coding1 benchmark
AA-SciCodeReported

Artificial Analysis SciCode

3.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Reasoning2 benchmarks
AA-LCRReported

Artificial Analysis Long Context Reasoning

0.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
CritPtReported

Critical Physics Tasks

0.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Knowledge6 benchmarks
Artificial Analysis Intelligence IndexReported
1.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

28.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

5.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

-83.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience AccuracyReported

Artificial Analysis Omniscience Accuracy

5.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience Hallucination RateReported

Artificial Analysis Omniscience Hallucination Rate

94.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Multimodal4 benchmarks
Liquid Extract JSON ValidityProvider exact

Liquid image-to-JSON extraction JSON validity

99.6%Display only
Source: Liquid AI: LFM2.5-VL-1.6B-Extract model cardProvenance: Provider exact
Liquid Extract F1Provider exact

Liquid image-to-JSON extraction schema consistency F1

99.6%Display only
Source: Liquid AI: LFM2.5-VL-1.6B-Extract model cardProvenance: Provider exact
Liquid Extract VLM JudgeProvider exact

Liquid image-to-JSON extraction VLM judge score

90.6%Display only
Source: Liquid AI: LFM2.5-VL-1.6B-Extract model cardProvenance: Provider exact
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

26.5%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Inst. Following1 benchmark
AA-IFBenchReported

Artificial Analysis IFBench

33.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.

LFM2.5-VL Extract Family

1-6b

Frequently Asked Questions

How does LFM2.5-VL-1.6B-Extract perform overall in AI benchmarks?

LFM2.5-VL-1.6B-Extract has 15 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.

Is LFM2.5-VL-1.6B-Extract good for knowledge and understanding?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract good for coding and programming?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in coding and programming, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract good for reasoning and logic?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract good for agentic tool use and computer tasks?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in agentic tool use and computer tasks, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract good for multimodal and grounded tasks?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract good for instruction following?

LFM2.5-VL-1.6B-Extract has visible benchmark coverage in instruction following, but BenchLM does not currently assign it a global category rank there.

Is LFM2.5-VL-1.6B-Extract open source?

Yes, LFM2.5-VL-1.6B-Extract is an open weight model created by LiquidAI, meaning it can be downloaded and run locally or fine-tuned for specific use cases.

Which sibling models are related to LFM2.5-VL-1.6B-Extract?

LFM2.5-VL-1.6B-Extract belongs to the LFM2.5-VL Extract family. Related variants on BenchLM include LFM2.5-VL-450M-Extract.

Does LFM2.5-VL-1.6B-Extract have full benchmark coverage on BenchLM?

Not yet. LFM2.5-VL-1.6B-Extract currently has 15 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of LFM2.5-VL-1.6B-Extract?

LFM2.5-VL-1.6B-Extract has a published context window of 128K, which determines how much text it can process in a single interaction.

Last updated: July 23, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.