大模型算法工程师(行程规划方向)(MJ031345)
Company 携程
Focus 技术 · 算法
Shanghai
October 28, 2025
【岗位职责】
1、参与构建旅游领域AI Agent系统,主导旅游垂类模型的训练与调优,解决多城市、多约束条件下的动态路线优化问题。
2、设计强化学习(RL)框架,结合PPO、GRPO等算法实现,在行程规划领域达到优于顶尖LLM或Agent能力。
3、构建端到端模型评测体系:设计多维度评估指标,研发大模型与传统运筹算法的融合架构。
4、开发个性化推荐能力,整合文本、POI、用户行为数据,推动生成式AI在旅游场景的落地应用。
【任职要求】
基础要求:
1、计算机科学、应用数学或运筹学硕士及以上学历,1年以上大模型全流程实战项目经验(数据构建→训练→评测→部署)。
2、深度掌握模型训练与评测技术:
3、精通大模型微调技术及分布式训练框架(DeepSpeed、Megatron)
4、具备强化学习实战经验,熟悉PPO、GRPO等算法在决策优化场景的应用
5、掌握模型评测方法论(自动指标+人工评估)及A/B测试设计
优先条件:
1、RL项目经验:有基于强化学习的决策优化系统(如路径规划)落地经验者优先
2、大模型训练经验:主导过模型的训练/微调/评测全流程者优先
Job Responsibilities
Participate in building the AI Agent system in the tourism field, lead the training and tuning of tourism vertical models, and solve dynamic route optimization problems under multi-city and multi-constraint conditions.
Design reinforcement learning (RL) frameworks, implement them with algorithms such as PPO and GRPO, and achieve performance superior to top-tier LLMs or Agents in the itinerary planning field.
Construct an end-to-end model evaluation system: design multi-dimensional evaluation metrics and develop a fusion architecture of large models and traditional operations research algorithms.
Develop personalized recommendation capabilities, integrate text, POI, and user behavior data, and promote the practical application of generative AI in tourism scenarios.
Qualifications
Basic Requirements:
Master's degree or above in Computer Science, Applied Mathematics, Operations Research or related majors, with more than 1 year of practical experience in the full process of large model projects (data construction → training → evaluation → deployment).
Have a deep grasp of model training and evaluation technologies.
Proficient in large model fine-tuning technologies and distributed training frameworks (DeepSpeed, Megatron).
Possess practical experience in reinforcement learning, and be familiar with the application of algorithms such as PPO and GRPO in decision optimization scenarios.
Master model evaluation methodologies (automatic metrics + manual evaluation) and A/B test design.
Preferred Qualifications:
RL project experience: Priority is given to candidates with practical experience in deploying decision optimization systems (such as path planning) based on reinforcement learning.
Large model training experience: Priority is given to candidates who have led the full process of model training/fine-tuning/evaluation.