rbrt's picture
Update README.md
e17d261 verified
|
raw
history blame
1.96 kB
metadata
license: mit
library_name: transformers
base_model:
  - deepseek-ai/DeepSeek-V3-0324
  - deepseek-ai/DeepSeek-R1
pipeline_tag: text-generation

DeepSeek-R1T-Chimera

TNG Logo


Model merge of DeepSeek-R1 and DeepSeek-V3 (0324)

An open weights model combining the intelligence of R1 with the token efficiency of V3.

Announcement on X | LinkedIn post | Try it on OpenRouter

Model Details

  • Architecture: DeepSeek-MoE Transformer-based language model
  • Combination Method: Merged model weights from DeepSeek-R1 and DeepSeek-V3 (0324)
  • Release Date: 2025-04-27

Contact

Citation

@misc{tng_technology_consulting_gmbh_2025,
    author       = { TNG Technology Consulting GmbH },
    title        = { DeepSeek-R1T-Chimera },
    year         = 2025,
    month        = {April},
    url          = { https://huggingface.co/tngtech/DeepSeek-R1T-Chimera },
    doi          = { 10.57967/hf/5330 },
    publisher    = { Hugging Face }
}