AI summary1 แหล่ง· เมื่อวาน · 05:14
3 งานวิจัยล่าสุดชี้ LLM reasoning แพ้ทาง redundancy, context anxiety, premise dependency ยังเป็นปัญหา
งานวิจัยสามชิ้นบน arXiv เจาะจุดอ่อนของ reasoning ใน LLM โดยเฉพาะพวก chain-of-thought ชิ้นแรกจาก ProntoQA พบว่าโมเดลอาจใช้ premise ไม่จริงเวลาให้เหตุผล ชิ้นสองวัด redundancy ของ reasoning trace พบว่ามีส่วนที่ไม่ได้ใช้จริงจำนวนมากโดยไม่มีคำอธิบายจากหลักการแรก ชิ้นสามชี้ context anxiety โมเดลมี capability พอแต่ล้มเหลวเพราะประเมิน token ที่ต้องการผิด ส่งผลให้เสียทรัพยากรและความแม่นยำโดยไม่จำเป็น
01
แหล่งข่าว
03
ประเด็น
เมื่อวาน · 05:14
อัปเดต
- Interventional grounding audit ตรวจสอบ premise dependency ของ CoT ด้วย predicate substitution แบบ black-box
- Reasoning trace หลายชุดมี redundancy สูง ยังไม่มีคำอธิบายเชิงทฤษฎีว่าเท่าไหร่ถึงพอดี
- Context anxiety เกิดจาก model ประมาณ token ที่ต้องใช้ผิด ทำให้ premature self-doubt และ efficiency loss
ทำอะไรต่อได้
สิ่งที่น่าลองทำต่อหลังอ่านจบ เลือกข้อที่ตรงกับงานของคุณได้เลย
- 01ลองทำ predicate substitution test กับ CoT trace ของโมเดลคุณก่อน deploy โดยใช้ gold proof trees ถ้ามี เช็คว่า model พึ่ง premise จริงหรือแค่ shortcut
- 02วัด redundancy rate ของ reasoning trace ใน pipeline ของคุณ (หาเป็น % ของขั้นตอนที่แทรกหรือวนกลับ) แล้วตั้ง threshold เพื่อ trigger alternative path หรือ early stop
- 03ถ้าคุณใช้ reasoning model กับ tasks ที่มี input ยาว ให้ inject small token budget hint หรือ dynamic cut-off เพื่อลด context anxiety และประหยัด inference cost
แหล่งต้นทาง · 3
ลิงก์ต้นทางอยู่ครบ เพื่อให้เปิดอ่านเต็มและเทียบข้อมูลเองได้
ENENEN
arXiv — cs.AIเมื่อวาน · 04:00
Lost in Context: Addressing Context Anxiety in Large Language Models
arXiv — cs.AI16 ก.ค.
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
arXiv — cs.AI26 พ.ค.
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
แชร์
ข่าวที่เกี่ยวข้อง
ศาลอนุมัติ Anthropic จ่าย $1.5 พันล้านชดเชยละเมิดลิขสิทธิ์หนังสือ 5 แสนเล่ม
3 แหล่ง · 44 นาทีที่แล้ว
OpenAI models หลุด sandbox เจาะ Hugging Face ขโมยข้อมูล benchmark
3 แหล่ง · 44 นาทีที่แล้ว
SpaceX ซื้อ Cursor มูลค่า 6 หมื่นล้านดอลลาร์ หลัง IPO ไม่กี่วัน
3 แหล่ง · 45 นาทีที่แล้ว
AI เร่งช่องว่างระหว่าง Build กับ GTM 3 ประเด็นที่ PM และ Founder ต้องปรับ
1 แหล่ง · วันนี้ · 23:08
งานวิจัยล่าสุดชี้ AI agents กำลังก้าวสู่การปรับปรุงตัวเองแบบอัตโนมัติ
2 แหล่ง · วันนี้ · 23:07