Lightnews — Scholar-powered news

Juan Rodriguez

@joanrod.bsky.social

24 followers 24 following 9 posts

AI Researcher. Working on Multimodal AI at ServiceNow, Mila
joanrod.github.io

Posts Replies Media Videos

Reposted by Juan Rodriguez

Alexandre Lacoste

@alex-lacoste.bsky.social

We’re really excited to release this large collaborative work for unifying web agent benchmarks under the same roof.

In this TMLR paper, we dive in-depth into #BrowserGym and #AgentLab. We also present some unexpected performances from Claude 3.5-Sonnet

December 12, 2024 at 5:55 PM

Reposted by Juan Rodriguez

Chris Pal

@chrisjpal.bsky.social

LLMs have a lot of potential for science, but scientists can be particularly sensitive to factuality, nuances, and hallucinations. The new ScholarQABench benchmark in this paper looks pretty useful for the community to monitor progress on LLMs for science. arxiv.org/html/2411.14199

November 25, 2024 at 1:20 AM

Juan Rodriguez

@joanrod.bsky.social

🎉 Excited to introduce BigDocs!
An open, transparent multimodal dataset designed for:
📄 Documents
🌐 Web content
🖥️ GUI understanding
👨‍💻 Code generation from images
We’re also launching BigDocs-Bench:
➡️ Document, Web, GUI Visual reasoning
➡️ Converting images into JSON, Markdown, LaTeX, SVG, and more!

December 10, 2024 at 6:34 PM

Add to Home Screen

Light up
your news

Add to Home Screen

Light upyour news

Sign in to Lightnews

Sign up to start reading

Connect Bluesky

Connect with Bluesky

Light up
your news