FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol | Summary
26 Mar 2026 | Paper Review Tool Use LLM Evaluation Model Context Protocol Synthetic DataContents
- Summary
- 1 Introduction
- 2 FinMCP-Bench: A Financial MCP Benchmark
- 3 Experimentation
- 4 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol.
- 2026-03-26 (arXiv)
- Zhu, Jie, Tian, Yimin, Li, Boyang, Wu, Kehao, Liang, Zhongzhi, Li, Junhui, Zhang, Xianyin, Guo, Lifan, Chen, Feng, Liu, Yong, et al.
- Qwen DianJin Team, Alibaba Cloud Computing, YINGMI Wealth Management, Soochow University
- Paper
- Project Page
Summary
- FinMCP-Bench evaluates LLM agents’ selection and organization of financial tools exposed through the Model Context Protocol (MCP). It contains 613 samples, 65 tools, 10 main scenarios, and 33 sub-scenarios, covering single-tool, multi-tool, and multi-turn tasks.
- The authors combine anonymized production logs with dependency-chain-based query synthesis and persona-based dialogue simulation, followed by automated validation and expert review. Evaluation measures tool precision, recall, F1, and exact agreement with reference tool groups; each dialogue turn is generated using gold conversation history rather than the model’s previous responses.
- Qwen3-235B-A22B-Thinking achieves the highest reported overall TF1 of 64.27% and EMR of 25.92%, but no model exceeds 10.62% multi-tool EMR or 4.10% multi-turn EMR. These results measure agreement with reference tool workflows, not financial advice quality, argument correctness, or autonomous end-to-end dialogue success.
1 Introduction
Financial agents must interpret user intent, retrieve information such as stock trends and fund holdings, and coordinate tool calls that depend on earlier results. The paper argues that existing financial evaluations typically address specific tasks without testing tool use, leaving these operational dependencies insufficiently evaluated.
- MCP supplies a standardized schema for tool invocation across servers; FinMCP-Bench evaluates financial workflows using this interface.
- Construction begins with 10K production interaction records across 33 scenarios. The introduction describes synthetic chains exceeding five steps and a final set of 10K curated traces, but does not explain how that set relates to the 613-sample benchmark specified in the detailed construction sections.
2 FinMCP-Bench: A Financial MCP Benchmark
FinMCP-Bench contains 145 single-tool samples, 249 multi-tool samples, and 219 multi-turn samples. Figure 1 maps its scenario coverage: Market Analysis & Research (MAR) is the largest category with 141 samples, followed by Investment Planning & Allocation (IPA) with 101.
- Single-tool tasks require one call in one conversational turn; multi-tool tasks require several sequential or parallel calls within one turn.
- Multi-turn tasks span several conversational turns, each potentially requiring one or more tools.
- The remaining categories are Financial Planning (28), Account & Service Management (47), Transaction Execution & Operation (96), Investor Education (71), Product & Strategy Consulting (52), Compliance & Legal Affairs (17), Platform Technical Support (31), and Other Consulting (30).
The scenario map extends coverage beyond market analysis to account operations, education, compliance, and technical support. Coverage is uneven: MAR has 141 instances, whereas CLA has 17, so the aggregate evaluation draws on different numbers of examples from each financial scenario.
2.1 Data Source
The source data consists of 10,000 historical logs from the XiaoGu AI assistant in the Qieman APP operated by Yingmi Fund. The assistant follows expert-defined SOPs and workflow-style tool invocation; the authors report that all logs undergo strict anonymization and disclosure procedures.
- A retained log must express a genuine financial need, resolve the problem through tool calls, and provide a satisfactory final response.
- Filtering yields 1,484 single-tool samples and 183 multi-tool samples. The single-tool pool is randomly divided into 700 samples for expert selection and 784 samples reserved for more complex sample synthesis.
- Expert review retains 145 of the 700 single-tool candidates. Section 2.1 reserves the other 784 for both multi-tool and multi-turn synthesis, although Section 2.3 does not detail their use in dialogue generation.
2.2 Chain-based Multi-tool Sample Construction
Tool Dependency Graph Construction. The chain-based method in Figure 2 starts with one node per tool and no edges, then processes the real multi-tool samples. Candidate dependencies are proposed between tools in immediately consecutive execution groups, and Qwen3-235B-2507 judges whether each dependency is reasonable in a financial scenario.
- For consecutive groups (t1, t2) and (t3), the candidates are t1 → t3 and t2 → t3; no dependency is inferred between the parallel tools t1 and t2.
- Validated candidates become directed edges. The resulting graph contains 65 nodes and 288 edges.
User Query Generation. For each scenario, the authors sample two tools that appear in real multi-tool samples associated with that scenario and select a connecting path in the dependency graph. Qwen3-235B-2507 generates a query aligned with this chain, using one randomly selected single-tool example per chain tool from the reserved pool.
- When several paths connect the sampled tools, one path is selected randomly.
- The synthetic query is therefore conditioned on scenario-associated tools, a dependency path, and examples of individual tool use.
Trajectory Generation. Qwen3-235B-2507, connected to Qieman’s MCP server, generates a tool-use trajectory for each synthetic query. Retained trajectories must preserve the sampled chain’s dependency relations, although additional tools are allowed.
- Of 1K generated query-trajectory pairs, 496 pass the initial construction filter.
- These 496 candidates are combined with 183 real multi-tool samples for expert review, which retains 249 multi-tool samples. The final real-versus-synthetic composition is not reported.
The three-stage diagram connects production tool sequences to validated dependency edges, sampled paths, synthetic queries, and generated trajectories. Single-tool examples constrain query generation, while execution through the MCP server and dependency-preserving filtering constrain the resulting workflows.
2.3 Role-Playing-based Multi-turn Sample Construction
Figure 3 illustrates multi-turn synthesis through role-playing. A planner selects a financial customer persona and a sub-scenario, generates a corresponding user goal, and assigns Qwen3-235B-2507 to simulate both the user and assistant.
- Personas are drawn from the profile pool in Zhu et al. (2025a), with attributes including age, gender, and income level.
- The assistant interacts with financial tools while the simulated user issues follow-up requests.
- The authors generate 500 candidate dialogues. Qwen3-235B-2507 checks whether all user queries are addressed, retaining 378; financial expert review then retains 219.
Persona and goal selection establish the simulated user’s financial context, while follow-up requests extend the interaction beyond a single query. The assistant calls Qieman MCP tools during this dialogue; the construction text specifies that Qwen3-235B-2507 plays both conversational roles.
2.4 Quality Control
Quality control combines automated execution checks with review by six financial domain experts and experienced developers. Each sample is independently assessed by two randomly assigned reviewers on a 5-point Likert scale and accepted only when both award at least four on every dimension.
- The automated validator checks that all tools execute successfully without errors.
- The five review dimensions are question relevance, tool-chain completeness, tool-chain logical consistency, answer reliability and traceability, and data freshness.
- Reviewer disagreements are resolved through discussion. The paper does not report inter-rater agreement statistics.
2.5 Dataset Analysis
Tool trajectories are represented as successive execution groups, with tools inside a group invoked in parallel and their internal order interchangeable. Difficulty is defined by tool-call count: Easy samples require up to 5 calls, Medium samples require more than 5 and up to 10, and Hard samples require more than 10; Table 1 reports 343, 245, and 25 samples, respectively.
- All 145 single-tool samples contain exactly one call in one step.
- Multi-tool samples have median/average call counts of 8 / 7.32 and median/average step counts of 5 / 5.72. Parallel calls occur in 73 of the 249 samples.
- Multi-turn samples have median/average call counts of 4 / 5.00 and median/average turn counts of 6 / 5.95. The 65 Qieman MCP tools have 2.6 arguments on average.
The statistics distinguish tool-call count from execution steps and dialogue turns. Multi-tool tasks average 7.32 calls over 5.72 steps, with 73 samples containing parallel calls; multi-turn tasks average 5.00 calls across 5.95 turns. Only 25 of the 613 samples belong to the Hard split.
3 Experimentation
The experiments compare six LLMs across task types, financial scenarios, and tool-count difficulty levels. Evaluation separates agreement on tool selection from the stricter requirement to reproduce the reference’s sequential and parallel organization.
3.1 Experimental Settings
Models. The evaluated models are Qwen3-4B-Thinking, Qwen3-30B-A3B-Thinking, Qwen3-235B-A22B-Thinking, DeepSeek-R1, GPT-OSS-20B, and Seed-OSS-36B. Inference. Each response is generated from the current user utterance and the gold conversation history, and the invoked tools are extracted from the generated responses.
- Single-tool and multi-tool instances are treated as one-turn conversations; multi-turn instances preserve their dialogue structure.
- For turn i, the model receives the reference utterance-response pairs from earlier turns, not its own earlier generated responses. This follows the conversation evaluation setup in Zhu et al. (2025a) and excludes accumulated errors from an autonomous dialogue rollout.
- The paper does not specify decoding parameters, repeated-run variability, or detailed inference budgets.
3.2 Evaluation Metrics
Tool Recall (TR). Reference and predicted tool sets are compared while ignoring dependency relations; recall is the intersection size divided by the reference-set size. Tool Precision (TP). Precision uses the same intersection divided by the predicted-set size. Tool F1 (TF1). Their harmonic mean is TF1 = 2×TR×TP/(TR+TP).
- Set-based comparison measures tool selection but does not preserve execution grouping or repeated invocation counts.
- The paper motivates tool-focused evaluation by noting that financial research and advisory questions are open-ended and lack standard final answers.
Exact Match Rate (EMR). A prediction counts as exact only when its tool organization matches the reference, ignoring internal ordering within each parallel group. EMR therefore tests execution structure more strictly than TP, TR, and TF1.
- The metric definitions do not explicitly evaluate argument values, numerical correctness, or the quality of the final financial response.
- Exact agreement with the reference organization does not establish whether a different workflow could also solve the user’s request.
3.3 Experimental Results
Table 2 reports the highest overall TF1 and EMR for Qwen3-235B-A22B-Thinking, with TP 60.22%, TR 68.90%, TF1 64.27%, and EMR 25.92%. Model size does not produce a uniform ranking: Qwen3-4B-Thinking has higher overall EMR than Qwen3-30B-A3B-Thinking, 18.82% versus 18.24%, but lower TF1, 50.08% versus 55.58%.
- The remaining overall TF1 / EMR results are DeepSeek-R1 49.88% / 18.08%, GPT-OSS-20B 32.62% / 4.43%, and Seed-OSS-36B 39.34% / 13.86%.
- Qwen3-4B-Thinking leads single-tool TF1 and EMR at 68.55% and 65.52%, whereas Qwen3-235B-A22B-Thinking leads multi-tool TF1 and EMR at 69.42% and 10.62%.
Multi-turn EMR ranges from 0.00% to 4.10%, with Qwen3-30B-A3B-Thinking scoring highest. Qwen3-4B-Thinking has the highest multi-turn TF1 at 47.65%, showing that tool-set agreement and exact structural agreement produce different rankings.
- DeepSeek-R1 and GPT-OSS-20B have multi-turn TR of 5.15% and 6.03%, respectively, with EMR of 0.00% for both.
- Every model has higher recall on single-tool than multi-tool tasks, while single-tool precision is lower. The authors attribute the precision pattern to unnecessary extra tool selection when only one tool is required.
- The gap between multi-tool TF1 and EMR shows that recovering many reference tools does not imply recovering the complete reference organization.
Tool-set agreement is substantially higher than exact execution-organization agreement on multi-tool tasks. Qwen3-235B-A22B-Thinking leads overall TF1 and EMR, but multi-tool EMR never exceeds 10.62% and multi-turn EMR never exceeds 4.10%. Qwen3-4B-Thinking leads single-tool TF1 and EMR.
3.4 Experimental Analysis
Scenario-wise Results. Figure 4 compares TF1 across the ten main scenarios. The authors describe Qwen3-30B-A3B-Thinking and Qwen3-235B-A22B-Thinking as a leading group with broadly balanced profiles, and GPT-OSS-20B as trailing across the scenario axes.
- The authors attribute wider gaps to scenarios requiring multi-tool planning and cross-source synthesis, with narrower gaps for simpler single-operation queries. They do not provide a separate analysis isolating those workflow properties.
- Exact per-scenario values, confidence intervals, and significance tests are not supplied, limiting precise comparisons of small differences.
- Figure 4 labels Seed-OSS-32B, whereas the experimental settings and Table 2 identify Seed-OSS-36B. The PDF does not resolve this discrepancy.
The radar chart compares scenario-dependent tool-set agreement, with substantial overlap among several model profiles and GPT-OSS-20B generally occupying the smallest profile. Exact scenario scores are not tabulated. The legend identifies Seed-OSS-32B, unlike Seed-OSS-36B in the experimental settings and results table.
Difficulty-wise Results. Figure 5 shows that TF1 does not decrease monotonically as the required number of tool calls rises. The authors interpret this through precision-recall balance: Easy tasks penalize over-calling, whereas larger reference tool sets can reward improved recall.
- Qwen3-30B-A3B-Thinking and Qwen3-235B-A22B-Thinking improve from Easy to Hard; GPT-OSS-20B shows a substantial increase from Easy to Medium/Hard but remains below the others. Figure 5 also labels the Seed series Seed-OSS-32B rather than Seed-OSS-36B.
- Higher TF1 on the Hard split does not establish higher exact workflow success, because TF1 ignores dependency structure.
- Only 25 samples belong to the Hard split, compared with 343 Easy and 245 Medium samples. The paper reports no uncertainty estimates for these comparisons.
Hard-split TF1 exceeds Easy-split TF1 for every displayed model, although some Medium scores fall below Easy scores. This is compatible with the authors’ precision-recall explanation, but does not demonstrate that complex workflows are easier to execute correctly: TF1 ignores dependencies, and the Hard split contains only 25 samples. The Seed series is labeled Seed-OSS-32B rather than Seed-OSS-36B.
4 Conclusion
The paper concludes that financial MCP tool use remains challenging, particularly for multi-tool dependencies and multi-turn interaction. Its contribution is a benchmark and comparative tool-use evaluation, rather than a new agent-training method or evidence of improved financial decision quality.
Appendix
- The supplied nine-page PDF contains no appendix; pages 8–9 conclude with references. The first page lists https://huggingface.co/DianJin, https://modelscope.cn/organization/tongyi_dianjin, https://github.com/aliyun/qwen-dianjin, and https://tongyi.aliyun.com/dianjin. It states that the work has been submitted to the IEEE for possible publication, but identifies no specific publication venue.
Brief Thoughts
The benchmark’s useful distinction is between selecting reference tools and reproducing their execution structure. Production-derived tools, expert-reviewed samples, and parallel-call-aware EMR make the multi-tool gap concrete: Qwen3-235B-A22B-Thinking reaches 69.42% TF1 but only 10.62% EMR.
The evidence concerns reference-workflow agreement within one financial platform, not general financial-agent reliability. Gold-history inference excludes accumulated dialogue errors, the metrics do not explicitly score arguments or final answers, and Qwen3-235B-2507 performs much of the synthetic generation and automated judgment without a reported generator-bias analysis. The final real-versus-synthetic composition of the 249 multi-tool samples is not reported, preventing a separate assessment of authentic production queries from the published results. The introduction’s statement about 10K curated traces is not reconciled with the detailed 613-sample benchmark construction. The absence of uncertainty estimates, the small Hard split, and the Seed-OSS labeling mismatch in both analysis figures limit interpretation of fine-grained rankings.