Skip to content
Main Site News Console

How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

· Hugging Face Translated
教程模型卡

Three months ago, we began the revival of Papers with Code (see also the announcement tweet). Its goal is to make open AI research accessible and understandable, enabling people to easily find artifacts related to papers, learn about the state of the art (SOTA) across different areas of AI, share interesting research, and continue building on one another’s work. In other words, its goal is to advance the research frontier and foster the next Transformer.

Of course, making AI research accessible requires a powerful search engine that allows both humans and agents to quickly find relevant and related work through the website or the pwc search CLI command; agents can use the command through a Skill.

It is important to note that searching research is not exactly the same as searching ordinary text. A useful paper search engine should be able to find an exact title or arXiv identifier, but it should also understand queries such as “small language models for code generation,” even when those words do not appear together in that form in a paper. It also needs to recognize navigational requests such as “the original BERT paper,” tolerate incomplete titles or typos, and return results quickly even when the model service is cold-starting or temporarily unavailable.

Papers with Code search results for the query DINO

Papers with Code search results for the query DINO.

For Papers with Code, we built a hybrid search system. This was also informed by our previous experience at ML6, where we developed RAG-based systems for clients. Hybrid search has generally proven superior to keyword- or vector-based search systems because it combines the strengths of both approaches (for more information, see this blog post). Keyword search can find exact matches, while vector search can identify fuzzier, semantically similar terms. It is worth noting that a reranker (also known as a cross-encoder) can further improve results, but it also introduces additional cost and latency.

Hybrid retrieval outperforms vector-only or keyword-only search

Hybrid retrieval outperforms vector-only or keyword-only search. Chart from Microsoft’s Azure AI Search: Outperforming Vector Search with Hybrid Retrieval and Reranking (2023).

Papers with Code relies on a PostgreSQL database, whose full-text search capabilities provide a fast lexical-search baseline. For dense embeddings, we use pgvector to improve semantic recall, and combine the two using the reciprocal rank fusion (RRF) algorithm. Dense embeddings are supported by three Hugging Face services:

At present, the system maintains embeddings for more than 110,000 existing papers from arXiv and Daily Papers. This article describes its architecture, the design decisions behind it, and what we learned while putting it into production.

TL;DR

We intentionally split search into two parts: offline corpus construction and an online search service:

Architecture diagram of the offline corpus construction and online hybrid search pipeline

Architecture of the offline corpus construction and online hybrid search pipeline.

Expensive, throughput-oriented work is handled by Jobs. Persistent artifacts are stored in a Bucket. Only the small step of embedding a query sits in the request path, supported by a protected Inference Endpoint to serve online search. If the endpoint is cold-starting, busy, or unhealthy, search immediately falls back to full-text search. This separation gives the system both powerful capabilities and high speed.

Start with a Strict Embedding Contract

Embedding pipelines often fail in subtle ways: the model version changes, query and document prompts are mixed up, vector truncation is inconsistent, or an updated abstract no longer matches the stored vector.

To avoid these issues, we treat the embedding format as a versioned API. Each paper is encoded in the following format:

normalized title + "\n\n" + normalized abstract

For every vector-generation run, we record:

  • The model repository and its exact revision;
  • The output dimensions;
  • The input-format version;
  • Whether the input is a query or a document;
  • The normalization method;
  • Content hashes of the source title and abstract.

Our production environment uses Qwen/Qwen3-Embedding-0.6B, pinned to an exact revision, and generates 256-dimensional, L2-normalized vectors. We selected this model with the help of the MTEB leaderboard. MTEB is a widely used benchmark for comparing embedding models. It is worth noting that newer embedding models such as Qwen3 support two new capabilities:

  • They allow you to specify a dynamic embedding size, enabling a trade-off between quality and speed and storage cost. Qwen models refer to this as “MRL,” or Matryoshka Representation Learning. You can learn more here. We chose an embedding dimension of 256 to improve search speed.

  • They allow you to provide instruction prompts. Qwen embedding models support a document prompt (which we use to embed papers), while online search uses its query prompt (to embed user queries).

This contract is maintained throughout the entire embedding process, from export through GPU inference and writing to PostgreSQL, to its eventual use for online retrieval.

Jobs Turn Database Snapshots into a Vector Corpus

Embedding the full corpus is a typical batch-processing task. It requires GPU access for a relatively short period, benefits from high throughput, and should not consume resources continuously between runs. Hugging Face Jobs is a great fit for this kind of task: a Job is defined by a command, hardware specifications, and an optional Docker image, and can run uv scripts with their dependencies declared inline.

Our corpus-construction pipeline first exports the latest version of each paper from a repeatable-read PostgreSQL snapshot. The exporter reads rows as a stream rather than loading the entire catalog into memory; it writes size-limited JSONL shards and creates a manifest containing row counts and SHA-256 checksums.

We synchronize this immutable run directory to a private Storage Bucket.

#教程#模型卡#开源#Omniapi.co

Published by the 4All API team

Original link:https://huggingface.co/blog/pwc-search

Need a mainstream LLM API? 4All API gives you one key to call OpenAI, Anthropic, Google Gemini, Qwen, DeepSeek, and dozens more — at official-pass-through pricing, with enterprise-grade reliability, integrated in 5 minutes.

Sign up for the 4All API console →