Ling-3.0-flash-Fin
Released 
August 2026


Ling-3.0-flash-Fin

Finance-tuned MoE model built for source-grounded investment research, valuation modeling, and long-horizon financial agents.

Modality:
Text
PDF
model ID
inclusionai/ling-3.0-flash-fin:free

Output Speed *

161.87
tok/s

Intelligence Index *

22.6
/ 100

Context Window *

262144
tokens

Input price

0
Anytoken

Output price

0
Anytoken
Ling 3.0 Flash Fin: A Finance-Specialized MoE for Source-Grounded Investment Workflows Ling 3.0 Flash Fin is a finance-enhanced variant of Ling 3.0 Flash from inclusionAI (Ant Group), developed with financial institutions and domain experts. It keeps the 124B-total, 5.1B-active Mixture-of-Experts architecture and 256K context window of the base model, adding continued training on financial data. It targets end-to-end investment research: source-grounded retrieval, multi-document reasoning across filings and earnings reports, valuation and spreadsheet modeling, and report preparation. It retains general reasoning, coding, and math ability, making it best suited to tool-intensive, long-horizon finance agents rather than open-domain chat. Access Ling 3.0 Flash Fin via the AnyAPI.ai API and add a finance-specialized layer to your stack.

Performance

Where the Finance Tuning Actually Shows Up

Ling 3.0 Flash Fin is built for tool-intensive financial work rather than raw benchmark leadership. inclusionAI positions its strengths in source selection, multi-document reconciliation across filings, and spreadsheet-heavy valuation tasks, and reports vendor results tying near-top scores on SpreadsheetBench V1. On the Artificial Analysis Intelligence Index it scores 23, above the median of 8 for its size class, and generates output at roughly 159 tokens per second. In production this means the model earns its place on structured, source-grounded finance pipelines where traceability and spreadsheet fidelity matter more than general chat quality.

Benchmarks

Independent Signals vs. Vendor Finance Claims

Independent coverage is still thin. Artificial Analysis measures Ling 3.0 Flash Fin at an Intelligence Index of 23 and around 159 tokens per second, noting it is faster than average but very verbose. inclusionAI's own finance results—across FinFIRST, FinSearchComp Verified, FinCRAFT, Finance Agent, APEX-Agents, SpreadsheetBench, and τ³-Banking—are vendor-reported, including strong SpreadsheetBench V1 numbers but weaker APEX-Agents and SpreadsheetBench V2 scores. Because FinFIRST was open-sourced under Apache 2.0, third parties can verify these claims, but no independent run has been published yet. Treat finance-specific figures as provisional.

Output Speed

*
161.87
tok/s

Intelligence Index

*
22.6
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
23
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Ling 3.0 Flash Fin is a text-in, text-out Mixture-of-Experts model with 124B total and 5.1B active parameters, a 262,144-token context window, and up to 32,768 output tokens. It supports hybrid instant/reasoning modes and function calling, but does not enforce structured JSON output via response_format. The large context is the key production advantage: it can hold full annual reports and multi-sheet workbooks in a single request. The capped 32K output, however, constrains very long generated reports, which may need chunking.
Verified Specifications — 
Ling-3.0-flash-Fin
*
Input modalities
Text
PDF
output modalities
Text
Context window
262144
 tokens
Maximum output tokens
32768
Reasoning
Yes
Knowledge cutoff
August 2026
Pricing (standard)
0
 AnyTokens in
 / 
0
 AnyTokens out

Limitations & Trade-offs

Where Ling-3.0-flash-Fin falls short

1
No enforced structured output. Ling 3.0 Flash Fin accepts tools and tool_choice for function calling but does not support response_format, so JSON output is not enforced at the API level. For finance pipelines that require strict, schema-validated payloads—automated spreadsheet updates, downstream calculators, or structured investment memos—you must add your own validation and retry logic. The base Ling 3.0 Flash offers response_format, so teams needing reliable JSON may prefer it or a schema-enforcing model.
2
Vendor-only finance benchmarks. The finance-specific results across FinFIRST, FinCRAFT, SpreadsheetBench, APEX-Agents, and τ³-Banking are reported by inclusionAI, in some cases as a chart image rather than a numeric table, and no independent third-party run has been published. Weaker vendor scores on APEX-Agents and SpreadsheetBench V2 indicate harder agentic and spreadsheet tasks remain a gap. Teams should validate the finance tuning on their own workflows—FinFIRST's open Apache 2.0 release makes this feasible—before trusting it for high-stakes analysis.
3
High verbosity. Artificial Analysis flags Ling 3.0 Flash Fin as very verbose, generating around 250M tokens on the Intelligence Index versus a median of 100M. Combined with a 32,768-token output cap, verbose reasoning can crowd out final-answer space in long financial reports and inflate token throughput on high-volume routes. Workloads needing concise, tightly formatted outputs may require explicit length and formatting constraints, or a less verbose model.
4
Not investment advice, needs review. inclusionAI explicitly states that key assumptions, valuation results, and investment conclusions require professional review and do not constitute investment advice; as its first finance-enhanced release, it still needs validation in complex, long-horizon workflows. Production deployments must keep a human analyst in the loop for any regulated or decision-critical output, rather than treating model conclusions as final.

Best-Fit Workloads

Where this model earns its place

01

Source-grounded financial research agents


Fin is tuned to prioritize authoritative sources and produce traceable, cited answers, with FinFIRST built specifically to test retrieval, sourcing, and calculation traceability. Combined with function calling, it fits investment research assistants that pull from filings and data tools and must show their evidence chain. The trade-off: finance benchmark evidence is currently vendor-reported, so validate on your own sourcing tasks.

02

Multi-document filing analysis


The 256K context window lets the model ingest full annual reports, earnings releases, and regulatory filings in one request, and the finance tuning targets reconciling reporting periods, definitions, and conflicting figures across documents. This suits due-diligence and comparison tools that must reason over several long documents at once. Note the 32K output cap when generating comprehensive cross-document summaries.

03

Valuation and spreadsheet modeling


Fin is trained to understand Excel formulas, cross-sheet dependencies, actual-versus-estimate updates, balance checks, and scenario analysis, with vendor results tying near-top scores on SpreadsheetBench V1. It fits tools that build or update editable DCF and valuation models. Weaker reported SpreadsheetBench V2 performance means harder spreadsheet tasks still warrant careful review and testing.

04

Long-horizon banking and finance workflows


The MoE architecture with only 5.1B active parameters gives efficient inference for multi-step, tool-calling agent loops in banking and credit tasks, evaluated on τ³-Banking and Finance Agent benchmarks. It suits agents that chain retrieval, calculation, modeling, and report preparation. Because it is a first-generation finance release, keep a human reviewer on regulated or high-stakes execution paths.

Pricing in anytokens via AnyAPI
Input
0
Output
0
Cache write
0
Cache read
0

Integration

Access Ling-3.0-flash-Fin via AnyAPI.ai

Access Ling-3.0-flash-Fin through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Ling-3.0-flash-Fin without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Ling-3.0-flash-Fin and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Ling-3.0-flash-Fin against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Ling-3.0-flash-Fin from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Ling-3.0-flash-Fin for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Ling 3.0 Flash Fin is a finance-enhanced version of inclusionAI's Ling 3.0 Flash, built for real-world investment workflows. It targets source-grounded financial search, multi-document reasoning across filings and earnings reports, valuation and spreadsheet modeling, and report preparation, while retaining general reasoning, coding, and math capabilities. It is best suited to tool-intensive, long-horizon finance agents rather than open-domain chat.

Ling 3.0 Flash Fin has a 262,144-token (256K) context window and supports up to 32,768 output tokens. The large context lets it ingest full annual reports, multiple regulatory filings, or multi-sheet workbooks in a single request, while the output cap constrains very long generated reports, which may need to be split into sections.

Yes for function calling—Ling 3.0 Flash Fin accepts tools and tool_choice, enabling agent workflows that call external data or calculation tools. However, it does not support response_format, so JSON output is not enforced at the API level. Applications needing strict, schema-validated JSON should add their own validation or pair it with a schema-enforcing model.

Yes. inclusionAI (Ant Group) open-sourced Ling 3.0 Flash Fin with BF16 weights under an MIT license, and released the FinFIRST financial-search benchmark under Apache 2.0. Because it shares the base Ling 3.0 Flash architecture, it runs on the same SGLang and vLLM runtimes, so teams can self-host or use hosted API access.

Ling 3.0 Flash Fin shares the base model's 124B-total, 5.1B-active MoE architecture, 256K context, and function calling, but adds continued training on financial data for source-grounded retrieval, filing analysis, and valuation modeling. The base Ling 3.0 Flash is faster in independent measurement and supports response_format for JSON. Choose Fin for finance-specific work; choose the base model for general throughput or JSON output.

* Benchmark data source: Artificial Analysis artificialanalysis.ai