mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Financial

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

How to generate per-query evaluation rubrics that encode institution-specific standards and fix information cutoffs without manual review overhead.

Source: arxiv.org
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Financial research agents need evaluation frameworks that reflect institution-specific standards and fix values as of an information cutoff. Fixed benchmarks with hand-written rubrics are expensive to extend and cannot encode proprietary evaluation criteria. FinAutoRubric solves this by letting experts write reusable guidance once, then generating per-query rubrics automatically through a multi-agent loop with code-enforced validation.

The Evaluation Problem

When you deploy a financial research agent, you need to know whether its answers meet your institution’s standards. Generic benchmarks do not help because:

  • Institution-specific standards: Your firm’s risk thresholds, data sources, and compliance requirements differ from public benchmarks.
  • Information cutoff: Financial data changes. A correct answer on March 15 may be wrong on March 16. Rubrics must fix the expected value as of a specific date.
  • Cost of manual rubrics: Writing a detailed rubric for every query does not scale. Expert time is the bottleneck.

Existing expert-reviewed finance benchmarks rely on fixed, per-item rubrics. Each rubric is written once for a specific query and cannot be reused. Extending the benchmark requires writing new rubrics from scratch.

Architecture: Expert Guidance + Agent Loop

FinAutoRubric separates reusable expert guidance from query-specific rubric generation. The architecture has three layers:

  1. Expert guidance layer: Analysts write evaluation criteria once as prompts and code-enforced rules. This guidance governs all agents.
  2. Task Bank: A library of reusable evaluation criteria that carry expert standards across tasks.
  3. Multi-agent generation loop: A writer agent researches expected values, a reviewer agent verifies them, and failures escalate to a human.

Generation Flow

# Simplified rubric generation loop
def generate_rubric(query, expert_guidance, task_bank):
    """
    Generate a rubric for a financial research query.
    
    Args:
        query: The research question (e.g., "What is AAPL's P/E ratio?")
        expert_guidance: Prompts and rules from analysts
        task_bank: Reusable criteria library
    
    Returns:
        Rubric with expected values and scoring criteria
    """
    # Writer agent researches expected values
    draft_rubric = writer_agent.generate(
        query=query,
        guidance=expert_guidance,
        task_bank=task_bank,
        information_cutoff="2026-09-28"
    )
    
    # Reviewer agent verifies each criterion
    review_result = reviewer_agent.validate(
        rubric=draft_rubric,
        guidance=expert_guidance,
        sources=draft_rubric.sources
    )
    
    # Escalate failures to human
    if review_result.confidence < THRESHOLD:
        return escalate_to_human(draft_rubric, review_result)
    
    return draft_rubric

The writer agent queries financial data sources, computes expected values, and drafts rubric criteria. The reviewer agent checks whether the expected values match the sources and whether the rubric follows expert guidance. If the reviewer’s confidence falls below a threshold, the system escalates to a human analyst.

Encoding Institution-Specific Standards

The expert guidance layer solves the problem of encoding proprietary standards without leaking them into model training data. Analysts write:

  • Prompts: Natural language instructions that describe evaluation priorities (e.g., “Prefer SEC filings over news articles for earnings data”).
  • Code-enforced rules: Validation logic that checks rubric structure, source quality, and temporal consistency.

These artifacts live in your infrastructure, not in the model. The agent reads them at runtime, so you can update standards without retraining.

Task Bank Structure

The Task Bank stores reusable criteria templates:

Criterion TypeExampleReusability Scope
Data source preference”Use Bloomberg for real-time prices”All price queries
Temporal constraint”Fix values as of market close”All time-sensitive queries
Calculation method”Use trailing twelve months for P/E”All ratio queries
Compliance check”Verify data is not material non-public”All queries

When generating a rubric, the writer agent pulls relevant criteria from the Task Bank and instantiates them for the specific query.

Information Cutoff Handling

Financial data changes constantly. A rubric must fix the expected value as of a specific date to prevent temporal drift. FinAutoRubric handles this in three ways:

  1. Explicit cutoff parameter: Every rubric generation call includes an information_cutoff timestamp.
  2. Source versioning: The writer agent records the timestamp of each data source it queries.
  3. Reviewer validation: The reviewer agent checks that all expected values are sourced from data available before the cutoff.

If the writer agent cannot find a source dated before the cutoff, it flags the criterion as unverifiable and escalates.

Validation Without Manual Review

The core challenge is validating that an auto-generated rubric matches expert judgment without requiring the expert to review every rubric. FinAutoRubric uses three mechanisms:

  1. Reviewer agent: A second LLM verifies that the rubric follows expert guidance and that expected values match sources.
  2. Code-enforced rules: Deterministic checks (e.g., “Every criterion must cite a source”) run before the reviewer agent.
  3. Confidence threshold: If the reviewer’s confidence score falls below a threshold, the rubric escalates to a human.

The paper reports that on three expert-authored finance benchmarks, FinAutoRubric’s rubrics track expert scoring as closely as the strongest evaluated generator. In-house analysts preferred the auto-generated rubrics in a blind review.

Failure Modes

Failure ModeCauseMitigation
Hallucinated expected valueWriter agent invents a number not in sourcesReviewer agent checks source citations; escalate if confidence is low
Stale dataWriter agent uses data after cutoffSource versioning + reviewer validation
Misaligned criteriaRubric does not match expert guidanceCode-enforced rules + Task Bank templates
Reviewer false positiveReviewer approves a bad rubricHuman spot-checks on escalated cases; log all rubrics for audit

The escalation mechanism is critical. If the reviewer agent’s confidence is below threshold, a human analyst reviews the rubric before it is used. This creates a feedback loop: analysts see which rubrics fail and update the expert guidance or Task Bank accordingly.

Deployment Shape

A production deployment needs:

  • Expert guidance repository: Version-controlled prompts and rules that analysts can update.
  • Task Bank database: Indexed library of reusable criteria, searchable by query type and asset class.
  • Rubric generation service: API that accepts a query and returns a rubric, with caching for repeated queries.
  • Escalation queue: Human-in-the-loop interface for reviewing low-confidence rubrics.
  • Audit log: Immutable record of all generated rubrics, sources, and confidence scores.

The paper’s released FinAutoRubric Benchmark includes 100 queries across 78 tasks and eight asset classes. Rubrics generated by an earlier model generation still leave headroom for a later one, which suggests that the framework is extensible as models improve.

Observability

You need to monitor:

  • Escalation rate: Percentage of rubrics that require human review. High rates indicate poor expert guidance or weak models.
  • Reviewer confidence distribution: Track the confidence scores over time. A shift toward lower confidence may signal data quality issues.
  • Source freshness: Measure the lag between information cutoff and source timestamps. Large lags indicate stale data.
  • Rubric reuse rate: How often criteria from the Task Bank are reused. Low reuse suggests the Task Bank needs more templates.

Log every rubric generation with full provenance: query, expert guidance version, Task Bank version, writer agent output, reviewer agent output, and final rubric. This audit trail is essential for compliance and debugging.

Technical Verdict

Use FinAutoRubric when:

  • You need to evaluate financial research agents against institution-specific standards.
  • You have expert analysts who can write reusable guidance but cannot manually review every query.
  • You need to fix expected values as of an information cutoff to prevent temporal drift.
  • You want to extend your evaluation benchmark without writing per-query rubrics from scratch.

Avoid it when:

  • Your evaluation criteria are simple enough that fixed benchmarks suffice.
  • You do not have expert analysts to write the initial guidance and Task Bank.
  • Your queries do not require temporal consistency (e.g., static knowledge questions).
  • You cannot tolerate the latency of a multi-agent generation loop (typically seconds per rubric).

The framework trades generation latency for extensibility. If you need to evaluate thousands of queries per second, pre-generate rubrics offline and cache them. If you need real-time evaluation, this approach will bottleneck on the writer and reviewer agents.

Tags

agentic-ai orchestration infrastructure

Primary Source

arxiv.org ↗