Skip to main content
Weaviate Docs (migrated from docs.weaviate.io) Docs

Search documentation

Type to search this documentation.

On this pageOverview

ANN Benchmark

This vector database benchmark is designed to measure and illustrate Weaviate's Approximate Nearest Neighbor (ANN) performance for a range of real-life use cases.

To make the most of this vector database benchmark, you can look at it from different perspectives:

  • The overall performance – Review the benchmark results to draw conclusions about what to expect from Weaviate in a production setting.
  • Expectation for your use case – Find the dataset closest to your production use case, and estimate Weaviate's expected performance for your use case.
  • Fine Tuning – If you don't get the results you expect. Find the optimal combinations of the configuration parameters (efConstruction, maxConnections and ef) to achieve the best results for your production configuration. (See HNSW Configuration Tips)

For each benchmark test, we set these HNSW parameters:

  • efConstruction - Controls the search quality at build time.
  • maxConnections - The number of outgoing edges a node can have in the HNSW graph.
  • ef - Controls the search quality at query time.

For each set of parameters, we've run 10,000 requests, and we measured the following metrics:

  • The Recall@10 and Recall@100 - by comparing Weaviate's results to the ground truths specified in each dataset.
  • Multi-threaded Queries per Second (QPS) - The overall throughput you can achieve with each configuration.
  • Individual Request Latency (mean) - The mean latency over all 10,000 requests.
  • P99 Latency - 99% of all requests (9,900 out of 10,000) have a latency that is lower than or equal to this number
  • Import time - Since varying build parameters has an effect on import time, the import time is also included.

By request, we mean: An unfiltered vector search across the entire dataset for the given test. All latency and throughput results represent the end-to-end time that your users would also experience. In particular, these means:

  • Each request time includes the network overhead for sending the results over the wire. In the test setup, the client and server machines were located in the same VPC.
  • Each request includes retrieving all the matched objects from disk. This is a significant difference from ann-benchmarks, where the embedded libraries only return the matched IDs.

This section contains datasets modeled after the ANN Benchmarks. Pick a dataset that is closest to your production workload:

Dataset Number of Objects Vector Dimensions Distance metric Use case
SIFT1M 1 M 128 l2-squared SIFT is a common ANN test dataset generated from image data
DBPedia OpenAI 1 M 1536 cosine DBPedia dataset from MTEB embedded with OpenAI's ada002 model
MSMARCO Snowflake 8.8 M 768 l2-squared MSMARCO dataset embedded with Snowflake's Arctic Embed M v1.5 model
Sphere DPR 10 M 768 dot Meta's Sphere dataset (10M sample)

These are the results for each dataset:

QPS vs Recall for DBPedia OpenAI

DBPedia OpenAI Benchmark results

efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
256161684.32121831.292.14286
128161683.28119271.322.24168
384161684.69118431.322.26401
256162489108161.462.42286
128162488.08105401.52.56168
384162489.18104861.52.65401
128321688.55104261.512.67216
256321689.7298011.62.65439
256163291.5695631.652.83286
128322492.2591671.722.89216
384163291.7489021.774.21401
128163290.6587961.794.06168
384321689.9287341.815.44620
256322493.4280241.974.07439
128323294.2680111.973.34216
384164894.4979121.993.3401
384322493.6478941.993.23620
256164894.3678312.023.95286
256323295.1874342.123.58439
128164893.5772792.174.97168
128166495.0270612.233.63168
256166495.7570042.263.77286
384323295.4669102.283.76620
384166495.9469072.283.81401
128324896.2466132.384.14216
256324897.0760722.594.3439
128169696.5156722.774.51168
256169697.2456392.84.43286
384324897.3455772.824.51620
384169697.454572.894.73401
128326497.1452772.985.73216
256326497.9651413.075.04439
1281612897.3247943.35.16168
2561612897.9747203.355.34286
384326498.1445573.476.61620
3841612898.1444823.515.96401
128329698.0444643.515.7216
256329698.6940343.916.19439
1283212898.5137494.26.84216
384329698.936224.346.72620
2563212899.0432604.837.89439
3843212899.2530115.227.94620
1281625698.4329645.38.52168
2561625698.9329575.338286
3841625699.0728715.498.12401
1283225699.1523346.7311216
1281638498.8522796.9410.55168
2561638499.2621857.2511.32286
3841638499.3621577.3210.8401
2563225699.5620137.8412.29439
3843225699.6918568.5312.78620
1281651299.0818458.5613.06168
2561651299.4217838.8713.09286
3841651299.5217399.1213.7401
1283238499.3917259.1515.57216
2563238499.73151710.4216.34439
1283251299.53139211.3118.65216
3843238499.81136311.6118.18620
2563251299.81122012.9320.23439
3843251299.8696916.140.21620
efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
256169692.435794.47.14286
256164892.435354.467.74286
384164892.5235144.477.48401
256162492.435024.58.04286
256166492.434904.528.25286
384169692.5234794.547.56401
384162492.5234584.579.22401
384166492.5234544.578.05401
384163292.5234264.618.29401
256163292.434124.629.81286
128169691.8233914.658.4168
128164891.8233294.749.07168
128166491.8232954.759.04168
128162491.8232384.8611.84168
2561612894.3431684.978.41286
3841612894.4831295.038.25401
128163291.8230645.148.28168
384161692.5229995.2615.3401
256161692.429835.3115.16286
128329695.1829455.348.78216
128324895.1829355.368.78216
128326495.1829085.419.56216
128161691.8228995.4315.77168
128323295.1828945.4310.61216
128321695.1828335.5513.92216
256326496.0828095.629.19439
128322495.1827985.6112.65216
256323296.0827715.6910.4439
256329696.0827705.699.82439
256324896.0827645.719.67439
256321696.0827585.7311.03439
256322496.0827085.8211.66439
1281612893.7926515.338.1168
1283212896.5126096.0310.24216
384326496.2725996.0610.23620
384323296.2725936.0910.03620
384329696.2725616.1610.85620
384324896.2725456.1812.8620
2563212897.2924796.359.94439
3843212897.523276.7910.77620
2561625697.6222976.8910.73286
3841625697.7822367.0711.42401
1281625697.1321847.2312.76168
1283225698.5618568.5113.92216
2561638498.617968.7913.99286
3841638498.7117858.8413.76401
1281638498.1717698.9115.82168
2563225699.0816709.4515.47439
3843225699.22156210.116.11620
1281651298.68150810.4817.79168
2561651299.04150210.5116.31286
3841651299.15146610.7816.79401
2563238499.51128212.3219.32439
384322496.27127012.4726.87620
3843238499.6121612.9620.16620
384321696.27120113.127.9620
2563251299.68105714.9123.59439
3843251299.76101015.6623.91620
1283238499.1196016.3543.94216
1283251299.3647532.9959.23216
How to read the results table
  • Choose the desired limit using the tab selector above the table.

    The limit describes how many objects are returned for a query. Different use cases require different levels of QPS and returned objects per query.

    For example, at 100 QPS and limit 100 (100 objects per query) 10,000 objects will be returned in total. At 1,000 QPS and limit 10 (10 objects per query), you will also receive 10,000 objects in total as each request contains fewer objects, but you can send more requests in the same timespan.

    Pick the value that matches your desired limit in production most closely.

  • Pick the desired configuration

    The first three columns represent the different input parameters to configure the HNSW index. These inputs lead to the results shown in columns four through six.

  • Recall/Throughput Trade-Off at a glance

    The highlighted columns (Recall, QPS) reflect the Recall/QPS trade-off. Generally, as the Recall improves, the throughput drops. Pick the row that represents a combination that satisfies your requirements. Since the benchmark is multi-threaded and running on a 30-core machine, the QPS/vCore columns shows the throughput per single CPU core. You can use this column to extrapolate what the throughput would be like on a machine of different size. See also this section below outlining what changes to expect when running on different hardware.

  • Latencies

    Besides the overall throughput, columns seven and eight show the latencies for individual requests. The Mean Latency columns shows the mean over all 10,000 test queries. The p99 Latency shows the maximum latency for the 99th-percentile of requests. In other words, 9,900 out of 10,000 queries will have a latency equal to or lower than the specified number. The difference between mean and p99 helps you get an impression how stable the request times are in a highly concurrent setup.

  • Import times

    Changing the configuration parameters can also have an effect on the time it takes to import the dataset. This is shown in the last column.

This is the recommended configuration for this dataset. It balances recall, latency, and throughput to give you a good overview of Weaviate's performance.

efConstructionmaxConnectionsefRecall@10QPS (Limit 10)Mean Latency (Limit 10)p99 Latency (Limit 10)
256169697.24%56392.80ms4.43ms

QPS vs Recall for SIFT1M

SIFT1M Benchmark results

efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
128162485.81183470.861.4955
256161680.32181420.871.5391
384162486.94176330.891.61129
128161679.29173930.913.9555
384161680.43173660.914129
256321686.46167430.941.74106
256162486.74164370.963.891
128321684.28163180.962.9860
384322492.281579111.73162
128163289.681564214.0355
128323292.98155151.021.7760
384163290.53154001.023.94129
384321687.02153931.023.94162
256163290.45153851.024.0491
128322489.92152891.034.0660
256164894.51148881.061.8391
256322491.68144151.093.94106
128164893.84142071.114.1155
128166495.95137611.142.0455
384164894.6137381.153.59129
384323295.07137001.153.97162
256323294.63136321.164.05106
384166496.54131811.22.07129
128324896.1128171.232.6660
256324897.27125661.262.33106
256166496.45125651.263.5491
128326497.52121361.32.3460
384324897.56119851.323.88162
384326498.58112541.42.47162
128169697.88109591.443.5455
256326498.35109401.443.13106
256169698.2108391.464.1691
384169698.26106171.493.38129
256161289999251.592.7391
128329698.7497081.623.5560
1281612898.7797021.624.155
3841612899.0493431.693.92129
384329699.490071.763.74162
256329699.2688361.793.89106
1283212899.2884701.863.9460
2563212899.679371.994.15106
3843212899.6875722.093.88162
1281625699.6569962.263.755
2561625699.7665952.394.3591
3841625699.7665772.413.99129
1283225699.8159582.654.2260
2563225699.8756112.824.64106
1281638499.8252413.025.3155
3843225699.8951693.065.21162
2561638499.8651203.085.3591
3841638499.8749503.196.2129
1283238499.8944663.556.2860
1281651299.8743643.636.1455
2561651299.9143093.685.6291
3841651299.9141353.836.26129
2563238499.9240063.957.64106
3843238499.9238864.068.09162
1283251299.9137184.266.4660
2563251299.9435264.496.81106
3843251299.9532894.818.04162
efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
128166491.855122.868.355
128164891.854542.889.0355
256169692.554472.898.6691
256164892.554302.98.0991
128169691.854142.98.8155
128163291.854122.918.9755
256161692.554082.918.4691
256163292.554002.929.0291
256166492.553352.948.6791
256162492.553302.959.5591
384166492.6253072.968.19129
128161691.852852.968.8655
384161692.6252352.998.31129
128326494.6752352.998.9760
128321694.67522938.1360
384164892.6252103.029.57129
128324894.6752023.028.5460
384162492.6251903.0211.44129
384163292.6251703.049.33129
128323294.6751593.059.4860
128329694.6751393.028.0760
384169692.6251053.079.23129
1281612894.3950983.078.9755
256323296.1550973.088106
256324896.1550913.098.05106
256329696.1550843.098.55106
256322496.1550693.18.31106
384326496.5850433.128.79162
384329696.5850423.118.34162
256326496.1550343.138.74106
384324896.5850293.138.43162
128162491.850233.1210.555
256161289550203.128.4891
384323296.5849883.149.25162
256321696.1549873.159.43106
1283212896.5249513.178.5960
3841612895.1249333.199.75129
128322494.6747943.279.4960
384321696.5847653.39.05162
2563212897.6447653.319106
384322496.5846983.359.38162
3843212897.9746543.388.39162
1281625698.3842343.728.9255
2561625698.6841023.839.8191
3841625698.7540753.869.36129
1283225699.1339553.979.4660
2563225699.5437784.178.81106
3843225699.6436824.278.56162
1281638499.3236784.298.5555
2561638499.4835664.429.0391
3841638499.5134454.5911.01129
1283238499.6533724.678.9660
2563238499.8431864.958.97106
1281651299.6431814.949.0955
3843238499.8831415.039.21162
2561651299.7431245.049.1391
3841651299.7631115.079.24129
1283251299.8229605.3410.3660
2563251299.9227825.6810.3106
3843251299.9426915.8510.48162
How to read the results table
  • Choose the desired limit using the tab selector above the table.

    The limit describes how many objects are returned for a query. Different use cases require different levels of QPS and returned objects per query.

    For example, at 100 QPS and limit 100 (100 objects per query) 10,000 objects will be returned in total. At 1,000 QPS and limit 10 (10 objects per query), you will also receive 10,000 objects in total as each request contains fewer objects, but you can send more requests in the same timespan.

    Pick the value that matches your desired limit in production most closely.

  • Pick the desired configuration

    The first three columns represent the different input parameters to configure the HNSW index. These inputs lead to the results shown in columns four through six.

  • Recall/Throughput Trade-Off at a glance

    The highlighted columns (Recall, QPS) reflect the Recall/QPS trade-off. Generally, as the Recall improves, the throughput drops. Pick the row that represents a combination that satisfies your requirements. Since the benchmark is multi-threaded and running on a 30-core machine, the QPS/vCore columns shows the throughput per single CPU core. You can use this column to extrapolate what the throughput would be like on a machine of different size. See also this section below outlining what changes to expect when running on different hardware.

  • Latencies

    Besides the overall throughput, columns seven and eight show the latencies for individual requests. The Mean Latency columns shows the mean over all 10,000 test queries. The p99 Latency shows the maximum latency for the 99th-percentile of requests. In other words, 9,900 out of 10,000 queries will have a latency equal to or lower than the specified number. The difference between mean and p99 helps you get an impression how stable the request times are in a highly concurrent setup.

  • Import times

    Changing the configuration parameters can also have an effect on the time it takes to import the dataset. This is shown in the last column.

This is the recommended configuration for this dataset. It balances recall, latency, and throughput to give you a good overview of Weaviate's performance.

efConstructionmaxConnectionsefRecall@10QPS (Limit 10)Mean Latency (Limit 10)p99 Latency (Limit 10)
256326498.35%109401.44ms3.13ms

QPS vs Recall for MSMARCO Snowflake

MSMARCO Snowflake Benchmark results

efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
128161685.01126341.252.391286
384161687.78121861.32.33350
256321691.24113321.42.442907
384321691.9109341.452.534366
256162491.04108441.462.712304
128163291.44107081.482.61286
128321688.9106691.484.51450
128322492.11105701.52.641450
384163293.58102161.552.713350
128323293.899691.592.791450
256163293.0898191.63.622304
256322494.0694211.683.672907
256323295.592011.722.992907
384162491.5591421.744.543350
384322494.6290281.763.844366
256161687.289681.774.352304
128164893.8288391.793.811286
384164895.6588181.83.13350
384323295.9787861.813.124366
128166495.0985411.863.211286
128324895.5981651.943.811450
256324896.9678062.033.522907
128326496.5474972.123.71450
384166496.774202.144.513350
256166496.3474102.144.522304
384324897.3673632.153.694366
128169696.4370392.253.821286
256326497.765182.444.672907
384169697.7763392.514.133350
256169697.4663372.514.382304
128329697.5260962.614.541450
1281612897.1758212.734.991286
256329698.4554292.934.942907
2561612898.0552403.036.032304
3841612898.3252193.045.723350
384329698.7150603.145.24366
1283212898.0450083.175.791450
2563212898.8245223.515.982907
3843212899.0442343.766.184366
1281625698.2938644.126.621286
2561625698.9234294.637.472304
3841625699.1333674.727.343350
1283225698.8632324.928.511450
1281638498.7228435.69.21286
2563225699.3528215.649.412907
3843225699.5125616.2110.514366
3841638499.3925016.369.683350
1283238499.1623716.7111.771450
1281651298.9323056.9110.811286
2561638499.2321497.4112.542304
2563238499.5320807.6612.752907
3841651299.5219718.0812.743350
2561651299.3919638.1113.652304
1283251299.3219198.2914.181450
3843238499.6619038.3613.634366
2563251299.6316399.7216.492907
3843251299.74149010.6917.914366
efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
128166492.0339004.067.521286
128169692.0338804.077.771286
128164892.0338554.118.31286
256164893.0838094.167.932304
128163292.0337924.179.141286
256162493.0837844.188.122304
256169693.0837814.197.832304
256166493.0837494.228.612304
256163293.0837344.248.522304
128162492.0337184.269.971286
1281612893.8636074.48.111286
384166493.3335844.428.273350
384164893.3335684.448.713350
128161692.0335394.4811.541286
384169693.3335244.498.733350
2561612894.8334874.558.412304
256161693.0834684.5712.382304
128323294.0534354.628.691450
128326494.0534184.638.541450
128324894.0533894.689.511450
128329694.0533834.689.031450
128322494.0533594.7210.391450
3841612895.0832844.838.973350
256329695.4131934.979.122907
256324895.4131804.988.892907
1283212895.51317059.071450
256323295.4131645.019.292907
256326495.4131345.069.852907
384326495.831125.19.244366
256321695.4131115.111.672907
128321694.0531065.113.711450
384324895.830905.139.854366
384329695.830535.199.754366
384323295.830495.1910.384366
256322495.4130395.2210.92907
2563212896.6728765.5210.092907
3843212897.0227955.6810.094366
1281625696.9426905.8810.521286
384163293.3326895.913.673350
2561625697.6625336.2611.132304
3841625697.8724026.6111.83350
1283225697.8822946.9212.81450
1281638497.9421847.2712.21286
384322495.820877.618.124366
2563225698.5920877.6112.822907
2561638498.5120497.7512.992304
3843225698.8219578.1214.384366
3841638498.6919488.1613.153350
1283238498.6118368.6515.151450
1281651298.4218288.6914.991286
2561651298.9216969.3815.592304
384162493.3316579.5519.533350
2563238499.1316259.7916.622907
3841651299.0716159.8416.313350
384161693.3316159.8321.113350
3843238499.3153910.3216.824366
1283251298.96153510.3417.471450
384321695.8138511.4625.174366
2563251299.38135511.7419.432907
3843251299.52126212.621.114366
How to read the results table
  • Choose the desired limit using the tab selector above the table.

    The limit describes how many objects are returned for a query. Different use cases require different levels of QPS and returned objects per query.

    For example, at 100 QPS and limit 100 (100 objects per query) 10,000 objects will be returned in total. At 1,000 QPS and limit 10 (10 objects per query), you will also receive 10,000 objects in total as each request contains fewer objects, but you can send more requests in the same timespan.

    Pick the value that matches your desired limit in production most closely.

  • Pick the desired configuration

    The first three columns represent the different input parameters to configure the HNSW index. These inputs lead to the results shown in columns four through six.

  • Recall/Throughput Trade-Off at a glance

    The highlighted columns (Recall, QPS) reflect the Recall/QPS trade-off. Generally, as the Recall improves, the throughput drops. Pick the row that represents a combination that satisfies your requirements. Since the benchmark is multi-threaded and running on a 30-core machine, the QPS/vCore columns shows the throughput per single CPU core. You can use this column to extrapolate what the throughput would be like on a machine of different size. See also this section below outlining what changes to expect when running on different hardware.

  • Latencies

    Besides the overall throughput, columns seven and eight show the latencies for individual requests. The Mean Latency columns shows the mean over all 10,000 test queries. The p99 Latency shows the maximum latency for the 99th-percentile of requests. In other words, 9,900 out of 10,000 queries will have a latency equal to or lower than the specified number. The difference between mean and p99 helps you get an impression how stable the request times are in a highly concurrent setup.

  • Import times

    Changing the configuration parameters can also have an effect on the time it takes to import the dataset. This is shown in the last column.

This is the recommended configuration for this dataset. It balances recall, latency, and throughput to give you a good overview of Weaviate's performance.

efConstructionmaxConnectionsefRecall@10QPS (Limit 10)Mean Latency (Limit 10)p99 Latency (Limit 10)
384324897.36%73632.15ms3.69ms

QPS vs Recall for Sphere DPR

Sphere DPR Benchmark results

efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
256161675.12106661.482.823464
128161672.44104681.52.881934
128321679.4893721.6832471
384162480.6191561.693.924964
256321682.8490951.733.094956
256163283.7490311.743.263464
256164887.3177232.043.573464
256322487.174652.075.774956
128163280.9874112.116.251934
384322487.6869902.216.717455
384323289.9964432.425.867455
384166489.8863332.484.484964
128324889.7762442.56.122471
384324892.9652852.976.87455
384169692.3750463.135.564964
128169689.749783.167.071934
128166487.0549023.29.121934
128326491.5547793.36.452471
1281612891.1841293.87.461934
3841612893.7940563.96.194964
128329693.4840383.897.42471
384329696.0635234.497.737455
2563212896.4430685.148.354956
1281625694.0626026.0120.881934
2561625695.8225846.129.523464
2561638496.7720277.8211.123464
3841638497.2519018.311.884964
1281651295.9316269.7161934
3841651297.79151410.4614.634964
3843238498.89121313.0520.437455
2563251298.86104515.1122.634956
efConstructionmaxConnectionsefRecallQPSMean Latencyp99 LatencyImport Time
12816328034234.67.251934
256162481.5633914.666.953464
256169681.5633654.697.113464
256166481.5633514.697.183464
256161681.5632784.817.463464
384164881.9432144.97.584964
384163281.9432084.927.724964
12816968028395.549.181934
12816168028335.5514.611934
128321685.0828305.578.922471
384161681.9427945.6214.524964
128324885.0827725.699.232471
128323285.0827295.759.292471
256164881.5627125.8116.753464
128329685.0826925.869.642471
256329687.4126915.858.854956
384326487.9926515.968.737455
384324887.9926475.978.857455
384169681.9425696.149.654964
256323287.4125586.169.524956
256324887.4124596.4114.434956
12816648024456.4419.651934
12816248023346.7119.781934
384166481.9422037.0619.854964
2563212890.0121787.2210.544956
256322487.4121757.2211.094956
384323287.9921747.2811.527455
2561612884.7121687.2520.033464
256163281.5621517.315.823464
128326485.0821387.3623.082471
1281612883.0721027.4421.981934
3841612885.120657.6411.854964
1283212887.7820197.7615.792471
384162481.9420037.8919.264964
3843212890.5819648.0612.987455
256326487.4118768.4325.454956
384329687.9917289.14257455
128322485.0816949.3225.242471
12816488016559.5218.121934
256321687.4116459.5924.744956
1281638492.55139211.3520.151934
3843225695.58137311.5116.247455
2561625691.31135011.732.293464
384322487.99132211.9730.257455
2561651295.36131311.9716.543464
3841638494.29130812.0621.334964
1283225693.33127512.430.152471
1281625689.76126612.4638.381934
2561638493.94125012.6819.43464
2563225695.08123412.8130.894956
3841625691.71123012.8633.164964
1283238495.44116013.528.572471
2563238496.86115513.7219.794956
384321687.99108714.5329.067455
3841651295.72107314.7737.294964
1281651294.07106514.7936.231934
1283251296.52101215.625.022471
2563251297.7696416.4323.954956
3843238497.2696316.4233.557455
3843251298.0991417.324.797455
How to read the results table
  • Choose the desired limit using the tab selector above the table.

    The limit describes how many objects are returned for a query. Different use cases require different levels of QPS and returned objects per query.

    For example, at 100 QPS and limit 100 (100 objects per query) 10,000 objects will be returned in total. At 1,000 QPS and limit 10 (10 objects per query), you will also receive 10,000 objects in total as each request contains fewer objects, but you can send more requests in the same timespan.

    Pick the value that matches your desired limit in production most closely.

  • Pick the desired configuration

    The first three columns represent the different input parameters to configure the HNSW index. These inputs lead to the results shown in columns four through six.

  • Recall/Throughput Trade-Off at a glance

    The highlighted columns (Recall, QPS) reflect the Recall/QPS trade-off. Generally, as the Recall improves, the throughput drops. Pick the row that represents a combination that satisfies your requirements. Since the benchmark is multi-threaded and running on a 30-core machine, the QPS/vCore columns shows the throughput per single CPU core. You can use this column to extrapolate what the throughput would be like on a machine of different size. See also this section below outlining what changes to expect when running on different hardware.

  • Latencies

    Besides the overall throughput, columns seven and eight show the latencies for individual requests. The Mean Latency columns shows the mean over all 10,000 test queries. The p99 Latency shows the maximum latency for the 99th-percentile of requests. In other words, 9,900 out of 10,000 queries will have a latency equal to or lower than the specified number. The difference between mean and p99 helps you get an impression how stable the request times are in a highly concurrent setup.

  • Import times

    Changing the configuration parameters can also have an effect on the time it takes to import the dataset. This is shown in the last column.

This is the recommended configuration for this dataset. It balances recall, latency, and throughput to give you a good overview of Weaviate's performance.

efConstructionmaxConnectionsefRecall@10QPS (Limit 10)Mean Latency (Limit 10)p99 Latency (Limit 10)
384329696.06%35234.49ms7.73ms

This benchmark is open source, so you can reproduce the results yourself.

Setup with Weaviate and benchmark machine

This benchmark test uses one GCP instances to run both Weaviate and the Benchmark scripts:

  • a n4-highmem-16 instance with 16 vCPU cores and 128 GB memory.

Based on your throughput requirements, it is very likely that you will run Weaviate on a considerably smaller or larger machine in production.

We have outlined in the Benchmark FAQs what you should expect when altering the configuration or setup parameters.

We modeled our dataset selection after ann-benchmarks. The same test queries are used to test speed, throughput, and recall. The provided ground truths are used to calculate the recall.

We use Weaviate's Golang client to import data. We use Go to measure the concurrent (multi-threaded) queries. Each language has its own performance characteristics. You may get different results if you use a different language to send your queries.

For maximum throughput, we recommend using the Go or Java client libraries.

The complete import and test scripts are available here.

How can I get the most performance for my use case?

Section titled “How can I get the most performance for my use case?”

If your use case is similar to one of the benchmark tests, use the recommended HNSW parameter configurations to start tuning.

For more instructions on how to tune your configuration for best performance, see HNSW Configuration Tips.

What is the difference between latency and throughput?

Section titled “What is the difference between latency and throughput?”

The latency refers to the time it takes to complete a single request. This is typically measured by taking a mean or percentile distribution of all requests. For example, a mean latency of 5ms means that a single request takes, on average, 5ms to complete. This does not say anything about how many queries can be answered in a given timeframe.

If Weaviate were single-threaded, the throughput per second would roughly equal to 1s divided by mean latency. For example, with a mean latency of 5ms, this would mean that 200 requests can be answered in a second.

However, in reality, you often don't have a single user sending one query after another. Instead, you have multiple users sending queries. This makes the querying side concurrent. Similarly, Weaviate can handle concurrent incoming requests. We can identify how many concurrent requests can be served by measuring the throughput.

We can take our single-thread calculation from before and multiply it with the number of server CPU cores. This will give us a rough estimate of what the server can handle concurrently. However, it would be best never to trust this calculation alone and continuously measure the actual throughput. This is because such scaling may not always be linear. For example, there may be synchronization mechanisms used to make concurrent access safe, such as locks. Not only do these mechanisms have a cost themselves, but if implemented incorrectly, they can also lead to congestion, which would further decrease the concurrent throughput. As a result, you cannot perform a single-threaded benchmark and extrapolate what the numbers would be like in a multi-threaded setting.

All throughput numbers ("QPS") outlined in this benchmark are actual multi-threaded measurements on a 30-core machine, not estimations.

The mean latency gives you an average value of all requests measured. This is a good indication of how long a user will have to wait on average for their request to be completed. Based on this mean value, you cannot make any promises to your users about wait times. 90 out of 100 users might see a considerably better time, but the remaining 10 might see a significantly worse time.

Percentile-based latencies are used to give a more precise indication. A 99th-percentile latency - or "p99 latency" for short - indicates the slowest request that 99% of requests experience. In other words, 99% of your users will experience a time equal to or better than the stated value. This is a much better guarantee than a mean value.

In production settings, requirements - as stated in support plans - are often a combination of throughput and a percentile latency. For example, the statement "3000 QPS at p95 latency of 20ms" conveys the following meaning.

  • 3000 requests need to be successfully completed per second
  • 95% of users must see a latency of 20ms or lower.
  • There is no assumption about the remaining 5% of users, implicitly tolerating that they will experience higher latencies than 20ms.

The higher the percentile (e.g. p99 over p95) the "safer" the quoted latency becomes. We have thus decided to use p99-latencies instead of p95-latencies in our measurements.

What happens if I run with fewer or more CPU cores than on the example test machine?

Section titled “What happens if I run with fewer or more CPU cores than on the example test machine?”

The benchmark outlines a QPS per core measurement. This can help you make a rough estimation of how the throughput would vary on smaller or larger machines. If you do not need the stated throughput, you can run with fewer CPU cores. If you need more throughput, you can run with more CPU cores.

Adding more CPUs reaches a point of diminishing returns because of synchronization mechanisms, disk, and memory bottlenecks. Beyond that point, you should scale horizontally instead of vertically. Horizontal scaling with replication is also available.

What are ef, efConstruction, and maxConnections?

Section titled “What are ef, efConstruction, and maxConnections?”

These parameters refer to the HNSW build and query parameters. They represent a trade-off between recall, latency & throughput, index size, and memory consumption. This trade-off is highlighted in the benchmark results.

I can't match the same latencies/throughput in my own setup. How can I debug this?

Section titled “I can't match the same latencies/throughput in my own setup. How can I debug this?”

If you are encountering other numbers in your own dataset, here are a couple of hints to look at:

  • What CPU architecture are you using? The benchmarks above were run on a GCP c2 CPU type, which is based on amd64 architecture. Weaviate also supports arm64 architecture, but not all optimizations are present. If your machine shows maximum CPU usage but you cannot achieve the same throughput, consider switching the CPU type to the one used in this benchmark.

  • Are you using an actual dataset or random vectors? HNSW is known to perform considerably worse with random vectors than with real-world datasets. This is due to the distribution of points in real-world datasets compared to randomly generated vectors. If you cannot achieve the performance (or recall) outlined above with random vectors, switch to an actual dataset.

  • Are your disks fast enough? While the ANN search itself is CPU-bound, the objects must be read from disk after the search has been completed. Weaviate uses memory-mapped files to speed this process up. However, if not enough memory is present or the operating system has allocated the cached pages elsewhere, a physical disk read needs to occur. If your disk is slow, it could then be that your benchmark is bottlenecked by those disks.

  • Are you using more than 2 million vectors? If yes, make sure to set the vector cache large enough for maximum performance.

Where can I find the scripts to run this benchmark myself?

Section titled “Where can I find the scripts to run this benchmark myself?”

The repository is located here.

Have a question or feedback? Here's how to reach us.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu