Salesforce
/

Llama-xLAM-2-70b-fc-r

@@ -22,9 +22,10 @@ library_name: transformers
 <p align="center">
   <a href="https://apigen-mt.github.io/">[Homepage]</a>  |
-  <a href="https://github.com/SalesforceAIResearch/xLAM">[Github]</a> |
-  <a href="https://blog.salesforceairesearch.com/large-action-model-ai-agent/">[Blog]</a>
 </p>
 <hr>
@@ -38,7 +39,7 @@ The new **xLAM-2** series, built on our most advanced data synthesis, processing
 We've also refined the **chat template** and **vLLM integration**, making it easier to build advanced AI agents. Compared to previous xLAM models, xLAM-2 offers superior performance and seamless deployment across applications.
 <p align="center">
-<img width="100%" alt="Model Performance Overview" src="img/model_board.png">
 <br>
 <small><i>Comparative performance of larger xLAM-2-fc-r models (8B-70B, trained with APIGen-MT data) against state-of-the-art baselines on function-calling (BFCL v3, as of date 04/02/2025) and agentic (τ-bench) capabilities.</i></small>
 </p>
@@ -46,10 +47,9 @@ We've also refined the **chat template** and **vLLM integration**, making it eas
 ## Table of Contents
 - [Model Series](#model-series)
-- [Benchmark Results](#benchmark-results)
 - [Usage](#usage)
   - [Basic Usage with Huggingface Chat Template](#basic-usage-with-huggingface-chat-template)
-- [License](#license)
 - [Citation](#citation)
 ## Model Series
@@ -132,26 +132,25 @@ And then interact with the model using your preferred method for querying a vLLM
-<!-- ## Benchmark Results
-Note: **Bold** and <u>Underline</u> results denote the best result and the second best result for Success Rate, respectively.
-### Berkeley Function-Calling Leaderboard (BFCL)
-![xlam-bfcl](media/xlam-bfcl.png)
-*Table 1: Performance comparison on BFCL-v2 leaderboard (cutoff date 09/03/2024). The rank is based on the overall accuracy, which is a weighted average of different evaluation categories. "FC" stands for function-calling mode in contrast to using a customized "prompt" to extract the function calls.* -->
 ## Benchmark Results
 ### Berkeley Function-Calling Leaderboard (BFCL v3)
 <p align="center">
-<img width="80%" alt="BFCL Results" src="img/bfcl-result.png">
 <br>
 <small><i>Performance comparison of different models on BFCL leaderboard. The rank is based on the overall accuracy, which is a weighted average of different evaluation categories. "FC" stands for function-calling mode in contrast to using a customized "prompt" to extract the function calls.</i></small>
 </p>
 ### τ-bench Benchmark
 <p align="center">
-<img width="100%" alt="Pass^k curves" src="img/pass_k_curves_retail_airline.png">
 <br>
 <small><i>Pass^k curves measuring the probability that all 5 independent trials succeed for a given task, averaged across all tasks for τ-retail (left) and τ-airline (right) domains. Higher values indicate better consistency of the models.</i></small>
 </p>
@@ -165,9 +164,29 @@ This release is for research purposes only in support of an academic paper. Our
 For all Llama relevant models, please also follow corresponding Llama license and terms. Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.
-<!-- ## Citation
-If you find this repo helpful, please consider to cite our papers:
 ```bibtex
 @article{zhang2024xlam,
@@ -181,8 +200,10 @@ If you find this repo helpful, please consider to cite our papers:
 ```bibtex
 @article{liu2024apigen,
   title={Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets},
-  author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and Lan, Tian and Kokane, Shirley and Tan, Juntao and Yao, Weiran and Liu, Zhiwei and Feng, Yihao and others},
-  journal={arXiv preprint arXiv:2406.18518},
   year={2024}
 }
 ```
@@ -194,5 +215,5 @@ If you find this repo helpful, please consider to cite our papers:
   journal={arXiv preprint arXiv:2402.15506},
   year={2024}
 }
-``` -->

 <p align="center">
+  <a href="https://arxiv.org/abs/2504.03601">[Paper]</a>  |
   <a href="https://apigen-mt.github.io/">[Homepage]</a>  |
+  <a href="https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k">[Dataset (Coming Soon)]</a> |
+  <a href="https://github.com/SalesforceAIResearch/xLAM">[Github]</a>
 </p>
 <hr>
 We've also refined the **chat template** and **vLLM integration**, making it easier to build advanced AI agents. Compared to previous xLAM models, xLAM-2 offers superior performance and seamless deployment across applications.
 <p align="center">
+<img width="100%" alt="Model Performance Overview" src="https://github.com/apigen-mt/apigen-mt.github.io/blob/main/img/model_board.png?raw=true">
 <br>
 <small><i>Comparative performance of larger xLAM-2-fc-r models (8B-70B, trained with APIGen-MT data) against state-of-the-art baselines on function-calling (BFCL v3, as of date 04/02/2025) and agentic (τ-bench) capabilities.</i></small>
 </p>
 ## Table of Contents
 - [Model Series](#model-series)
 - [Usage](#usage)
   - [Basic Usage with Huggingface Chat Template](#basic-usage-with-huggingface-chat-template)
+- [Benchmark Results](#benchmark-results)
 - [Citation](#citation)
 ## Model Series
 ## Benchmark Results
 ### Berkeley Function-Calling Leaderboard (BFCL v3)
 <p align="center">
+<img width="80%" alt="BFCL Results" src="https://github.com/apigen-mt/apigen-mt.github.io/blob/main/img/bfcl-result.png?raw=true">
 <br>
 <small><i>Performance comparison of different models on BFCL leaderboard. The rank is based on the overall accuracy, which is a weighted average of different evaluation categories. "FC" stands for function-calling mode in contrast to using a customized "prompt" to extract the function calls.</i></small>
 </p>
 ### τ-bench Benchmark
+<p align="center">
+<img width="80%" alt="Tau-bench Results" src="https://github.com/apigen-mt/apigen-mt.github.io/blob/main/img/taubench-result.png?raw=true">
+<br>
+<small><i>Success Rate (pass@1) on τ-bench benchmark averaged across at least 5 trials. Our xLAM-2-70b-fc-r model achieves an overall success rate of 56.2% on τ-bench, significantly outperforming the base Llama 3.1 70B Instruct model (38.2%) and other open-source models like DeepSeek v3 (40.6%). Notably, our best model even outperforms proprietary models such as GPT-4o (52.9%) and approaches the performance of more recent models like Claude 3.5 Sonnet (new) (60.1%).</i></small>
+</p>
 <p align="center">
+<img width="80%" alt="Pass^k curves" src="https://github.com/apigen-mt/apigen-mt.github.io/blob/main/img/pass_k_curves_retail_airline.png?raw=true">
 <br>
 <small><i>Pass^k curves measuring the probability that all 5 independent trials succeed for a given task, averaged across all tasks for τ-retail (left) and τ-airline (right) domains. Higher values indicate better consistency of the models.</i></small>
 </p>
 For all Llama relevant models, please also follow corresponding Llama license and terms. Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.
+## Citation
+If you use our model or dataset in your work, please cite our paper:
+```bibtex
+@article{prabhakar2025apigenmt,
+  title={APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay},
+  author={Prabhakar, Akshara and Liu, Zuxin and Yao, Weiran and Zhang, Jianguo and Zhu, Ming and Wang, Shiyu and Liu, Zhiwei and Awalgaonkar, Tulika and Chen, Haolin and Hoang, Thai and Niebles, Juan Carlos and Heinecke, Shelby and Wang, Huan and Savarese, Silvio and Xiong, Caiming},
+  journal={arXiv preprint arXiv:2504.03601},
+  year={2025}
+}
+```
+Additionally, please check our other related works regarding xLAM and consider citing them as well:
+```bibtex
+@article{zhang2025actionstudio,
+  title={ActionStudio: A Lightweight Framework for Data and Training of Action Models},
+  author={Zhang, Jianguo and Hoang, Thai and Zhu, Ming and Liu, Zuxin and Wang, Shiyu and Awalgaonkar, Tulika and Prabhakar, Akshara and Chen, Haolin and Yao, Weiran and Liu, Zhiwei and others},
+  journal={arXiv preprint arXiv:2503.22673},
+  year={2025}
+}
+```
 ```bibtex
 @article{zhang2024xlam,
 ```bibtex
 @article{liu2024apigen,
   title={Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets},
+  author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and Lan, Tian and Tan, Juntao and Yao, Weiran and Liu, Zhiwei and Feng, Yihao and RN, Rithesh and others},
+  journal={Advances in Neural Information Processing Systems},
+  volume={37},
+  pages={54463--54482},
   year={2024}
 }
 ```
   journal={arXiv preprint arXiv:2402.15506},
   year={2024}
 }
+```