iFANN
    ค้นหาใน iFANN...
    เข้าสู่ระบบ
    หน้าแรก
    ข่าว
    วิดีโอ
    รูปภาพ
    GIF
    สำรวจ
    โพล
    รางวัล
    iFAMOUS
    วิกิ
    อนิเมะ
    ห้อง
    การแจ้งเตือน
    ข้อความ
    ที่บันทึกไว้
    โปรไฟล์
    วิกิรางวัลiFAMOUSอันดับอุตสาหกรรมรางวัลครีเอเตอร์รางวัลผู้ใช้ข้อกำหนดความเป็นส่วนตัวหลักเกณฑ์ชุมชนแจ้งลบ / DMCAช่วยเหลือนักพัฒนา

    © 2026 iFANN

    หน้าแรก
    ค้นหา
    ข้อความ
    การแจ้งเตือน
    โปรไฟล์

    โพสต์

    Nate
    Nate@nate_512
    💭Tech💭AI

    CommerceAgentBench Leaderboard

    There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    2w

    22 ถูกใจ1 ไม่ถูกใจ8 รีโพสต์1 ความคิดเห็น
    ?

    ความคิดเห็น

    ยังไม่มีความคิดเห็น มาเป็นคนแรกกันเถอะ!

    โพสต์

    Nate
    Nate@nate_512
    💭Tech💭AI

    CommerceAgentBench Leaderboard

    There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    2w

    22 ถูกใจ1 ไม่ถูกใจ8 รีโพสต์1 ความคิดเห็น
    ?

    ความคิดเห็น

    ยังไม่มีความคิดเห็น มาเป็นคนแรกกันเถอะ!