HAKARI-Bench

NanoMTEB-BR

Overview#

NanoMTEB-BR is the compact Brazilian Portuguese retrieval group derived from the six native Retrieval tasks in MTEB-BR. It covers Brazilian tax law, higher-education and central-bank FAQs, federal audit-court case law, medical question answering, and general passage retrieval. Unlike translated benchmark copies, all six tasks originate from Brazilian Portuguese sources. The group therefore tests whether a retrieval model handles local legal and institutional language, domain terminology, question-to-answer matching, and semantically related passages across corpora that range from a few hundred documents to 10,000-document hard-negative pools.

What This Group Measures#

The six tasks expose different relevance relations within one language. FAQ and medical tasks retrieve answer-bearing text for natural questions. The legal tasks retrieve statutes, rulings, or case-law material whose relevant passages may share specialized terminology without repeating the question. Quati tests general passage retrieval with native Brazilian Portuguese queries and hard negatives.

This mix makes the group useful for separating broad Portuguese semantic quality from domain-specific retrieval behavior. A model can perform well on short institutional FAQs while struggling with long or terminology-heavy legal documents, or it can retrieve general passages well without ranking concise medical answers accurately.

Task Families#

  • Tax-law retrieval: BRTaxQAR maps Brazilian tax questions to relevant statutes, regulations, administrative rulings, and related legal sources.
  • Higher-education FAQ retrieval: FaQuADIR retrieves answers about Brazilian higher-education institutions and services.
  • Central-bank FAQ retrieval: FaqBacenRetrieval retrieves public information from Banco Central do Brasil.
  • Public-sector case-law retrieval: JurisTCU retrieves Tribunal de Contas da UniĆ£o rulings from a hard-negative legal corpus.
  • Medical QA retrieval: MedPTRetrieval retrieves Brazilian Portuguese answers for health and medical questions.
  • General passage retrieval: Quati retrieves relevant native Brazilian Portuguese passages from a hard-negative corpus.

Dataset Shape#

TaskDomainQueriesDocumentsPositive qrels
BRTaxQARBrazilian tax law200461546
FaQuADIRHigher-education FAQ200243201
FaqBacenRetrievalCentral-bank FAQ2001,654201
JurisTCUFederal audit-court case law15010,0001,711
MedPTRetrievalMedical question answering200500204
QuatiGeneral passage retrieval4910,000892

All tasks use Brazilian Portuguese (pt) queries and documents. The FAQ and medical tasks have relatively small document pools and are close to single-positive retrieval. JurisTCU and Quati use larger hard-negative pools and many relevant passages for some queries, so their scores should not be interpreted as the same retrieval problem at a different scale.

Evaluation and Interpretation#

NanoMTEB-BR is a dedicated viewer scope and contributes six tasks to a complete HAKARI-Bench evaluation. Like other language-focused NanoMTEB suites, it is not an additional canonical Overall component; the six result files are still evaluated, stored, and available for direct comparison.

The initial validation compared 29 models that had complete results on both the official MTEB-BR Retrieval suite and NanoMTEB-BR. Ranking the models across the same six tasks produced a Spearman correlation of 0.9860 and a Kendall correlation of 0.9320. Twenty-eight of the 29 models remained within two Borda rank positions. This supports using NanoMTEB-BR as a close rank-preserving proxy for model comparison, while absolute scores remain specific to the smaller Nano corpora and query samples.

Training and Leakage Notes#

Useful training data includes non-overlapping Brazilian Portuguese FAQ pairs, medical QA, legal question-to-passage examples, public-sector search logs, and general passage-retrieval supervision. Hard negatives should come from the same institution or domain and share terminology while answering a different question.

Training and validation pipelines should exclude the Nano evaluation queries, qrels, positive documents, and overlapping upstream test records. Legal and institutional sources can contain repeated templates or near-duplicate text, so document-level deduplication alone may not prevent leakage.

Public Sources#

Metadata Summary#

FieldValue
Task pages6
Queries999
Split-local documents22,858
Positive qrels3,755
Languagespt
Categoriesnatural_language
Positives / query avg3.76

Task Metadata Summary#

TaskBacking datasetLangCategoryQueriesDocsPositivesBM25 nDCG@10Dense nDCG@10Reranking hybrid nDCG@10Best profile
BRTaxQARNanoMTEB-BRptnatural_language2004615460.46500.28420.3810BM25
FaqBacenRetrievalNanoMTEB-BRptnatural_language2001,6542010.47170.65770.5858Dense
FaQuADIRNanoMTEB-BRptnatural_language2002432010.88700.84250.8779BM25
JurisTCUNanoMTEB-BRptnatural_language15010,0001,7110.46880.52270.5490Reranking hybrid
MedPTRetrievalNanoMTEB-BRptnatural_language2005002040.43280.69740.5975Dense
QuatiNanoMTEB-BRptnatural_language4910,0008920.38150.64940.5202Dense