Researchers release MioFFAn, an open-source framework for annotating and formalizing scientific formulas with LLM-assisted workflows
The tool targets a documented scarcity of high-quality datasets for translating mathematical expressions into executable symbolic code, introducing modular automation and human-in-the-loop evaluation.
1 source · cross-referenced
- MioFFAn is an open-source, document-centric framework designed to accelerate annotation for formula formalization, a task hindered by a scarcity of high-quality, domain-specific datasets.
- The framework builds on the MioGatto architecture and adds features such as equation selection and aided symbolic code specification to support diverse scientific fields.
- MioFFAn incorporates partial automation via LLMs through modular sub-tasks with strict output formats, enabling iterative refinement and evaluation using standard NLP metrics.
- Authors report a preliminary evaluation demonstrating the efficacy of the human-in-the-loop approach in improving annotation workflows.
Researchers from Universitat Pompeu Fabra and the International Center for Numerical Methods in Engineering introduced MioFFAn, an open-source framework aimed at addressing a documented scarcity of high-quality, ground-truth datasets for formula formalization—the task of translating mathematical expressions in scientific literature into executable symbolic code.
Built on the MioGatto architecture, MioFFAn extends prior features to overcome structural limitations and introduces domain-specific functionalities such as selection of equations of interest and aided symbolic code specification. The framework allows users to configure custom taxonomies and properties for identified symbols and compatible symbolic operators, making it adaptable to specialized scientific fields.
MioFFAn is designed to integrate partial automation via large language models through a modular set of automated sub-tasks with strict output formats. This design enables researchers to iteratively refine automation strategies and compare approaches using standard NLP metrics.
The authors describe a preliminary evaluation that demonstrates the efficacy of the human-in-the-loop approach, positioning MioFFAn as a practical tool for accelerating annotation workflows in technical domains.
- Jul 28, 2026 · arXiv cs.CL
Researchers release GAND, a benchmark to study gender bias in machine translation through gender-ambiguous natural data
Trust79 - Jul 28, 2026 · arXiv cs.CL
Official conference reviewer guidelines outperform LLM-generated reviewer-imitating guidelines in automated peer review study
Trust79 - Jul 28, 2026 · arXiv cs.AI
Researchers propose SeT-Diff, a diffusion-based foundational model for HPC telemetry and time-series
Trust79