AI summary1 แหล่ง· เมื่อวาน · 05:06
arXiv เผย 7 งานวิจัย multimodal AI ใหม่ ตั้งแต่โมเดล AV-JEPA ไปจนถึง edge deployment
arXiv มีงานวิจัย multimodal AI ออกมาหลายชิ้นที่น่าสนใจสำหรับ dev ที่ทำงานด้าน audio-visual และ edge deployment โดยเฉพาะ AV-JEPA ที่ขยาย LeJEPA มาเป็น multimodal self-supervised learning ด้วย early-fusion Vision Transformer และ modality dropout ทำให้ architecture เรียบ ไม่ต้องใช้ decoder หรือ contrastive loss อีกชิ้นคือ OmniMem ที่จัดการ memory สำหรับ streaming audio-visual LLMs โดยแยกจัดการ visual และ audio context แก้ปัญหา token imbalance นอกจากนี้ยังมีงานวิจัยเกี่ยวกับการทำนาย latency สำหรับ LLMs บน edge devices ที่ช่วยเลือกโมเดลให้เหมาะกับ hardware แต่ละตัว ซึ่งเป็นประโยชน์โดยตรงสำหรับการ deploy จริง
01
แหล่งข่าว
03
ประเด็น
เมื่อวาน · 05:06
อัปเดต
- AV-JEPA ใช้ early-fusion และ modality dropout เพื่อ cross-modal alignment โดยไม่ต้องใช้ decoder หรือ contrastive loss
- OmniMem เสนอ modality-aware memory compression สำหรับ streaming audio-visual LLMs แยกจัดการ visual และ audio context
- มีงานวิจัยการทำนาย latency สำหรับ LLMs บน edge devices ที่ช่วยเลือกโมเดลตาม hardware และ runtime backend
ทำอะไรต่อได้
สิ่งที่น่าลองทำต่อหลังอ่านจบ เลือกข้อที่ตรงกับงานของคุณได้เลย
- 01ลองทดสอบ AV-JEPA backbone กับ multimodal dataset ของทีม โดยเฉพาะถ้าต้องการ architecture ที่ clean และ efficient โดยเทียบกับ baseline ที่ใช้ contrastive learning
- 02เปรียบเทียบ performance ของ OmniMem กับวิธีการบีบอัด token แบบอื่นใน long-video understanding task ก่อนเลือกใช้ใน production
- 03ใช้ framework การทำนาย latency จากงานวิจัยนี้เพื่อประเมิน LLM หลายตัวบน edge device ของคุณก่อน deploy จริง โดยวัดทั้ง prefill และ decode phase
แหล่งต้นทาง · 7
ลิงก์ต้นทางอยู่ครบ เพื่อให้เปิดอ่านเต็มและเทียบข้อมูลเองได้
ENENENENENENEN
arXiv — cs.AIเมื่อวาน · 04:00
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
arXiv — cs.AI20 ก.ค.
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
arXiv — cs.AI10 มิ.ย.
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
arXiv — cs.AI9 มิ.ย.
OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs
arXiv — cs.AI8 มิ.ย.
Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization
arXiv — cs.AI23 พ.ค.
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
arXiv — cs.AI17 เม.ย.
Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference
แชร์
ข่าวที่เกี่ยวข้อง
ศาลอนุมัติ Anthropic จ่าย $1.5 พันล้านชดเชยละเมิดลิขสิทธิ์หนังสือ 5 แสนเล่ม
3 แหล่ง · วันนี้ · 05:02
OpenAI models หลุด sandbox เจาะ Hugging Face ขโมยข้อมูล benchmark
3 แหล่ง · วันนี้ · 05:02
SpaceX ซื้อ Cursor มูลค่า 6 หมื่นล้านดอลลาร์ หลัง IPO ไม่กี่วัน
3 แหล่ง · วันนี้ · 05:02
AI เร่งช่องว่างระหว่าง Build กับ GTM 3 ประเด็นที่ PM และ Founder ต้องปรับ
1 แหล่ง · วันนี้ · 23:08
งานวิจัยล่าสุดชี้ AI agents กำลังก้าวสู่การปรับปรุงตัวเองแบบอัตโนมัติ
2 แหล่ง · วันนี้ · 23:07