Submitted by Gabriel Tomitsuka 24 Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows TextQL 3 2