You can not select more than 25 topics
Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
Hui Xiao
49623f9c8e
Account memory of big memory users in BlockBasedTable in global memory limit (#9748)
Summary:
**Context:**
Through heap profiling, we discovered that `BlockBasedTableReader` objects can accumulate and lead to high memory usage (e.g, `max_open_file = -1`). These memories are currently not saved, not tracked, not constrained and not cache evict-able. As a first step to improve this, similar to https://github.com/facebook/rocksdb/pull/8428, this PR is to track an estimate of `BlockBasedTableReader` object's memory in block cache and fail future creation if the memory usage exceeds the available space of cache at the time of creation.
**Summary:**
- Approximate big memory users (`BlockBasedTable::Rep` and `TableProperties` )' memory usage in addition to the existing estimated ones (filter block/index block/un-compression dictionary)
- Charge all of these memory usages to block cache on `BlockBasedTable::Open()` and release them on `~BlockBasedTable()` as there is no memory usage fluctuation of concern in between
- Refactor on CacheReservationManager (and its call-sites) to add concurrent support for BlockBasedTable used in this PR.
Pull Request resolved: https://github.com/facebook/rocksdb/pull/9748
Test Plan:
- New unit tests
- db bench: `OpenDb` : **-0.52% in ms**
- Setup `./db_bench -benchmarks=fillseq -db=/dev/shm/testdb -disable_auto_compactions=1 -write_buffer_size=1048576`
- Repeated run with pre-change w/o feature and post-change with feature, benchmark `OpenDb`: `./db_bench -benchmarks=readrandom -use_existing_db=1 -db=/dev/shm/testdb -reserve_table_reader_memory=true (remove this when running w/o feature) -file_opening_threads=3 -open_files=-1 -report_open_timing=true| egrep 'OpenDb:'`
#-run | (feature-off) avg milliseconds | std milliseconds | (feature-on) avg milliseconds | std milliseconds | change (%)
-- | -- | -- | -- | -- | --
10 | 11.4018 | 5.95173 | 9.47788 | 1.57538 | -16.87382694
20 | 9.23746 | 0.841053 | 9.32377 | 1.14074 | 0.9343477536
40 | 9.0876 | 0.671129 | 9.35053 | 1.11713 | 2.893283155
80 | 9.72514 | 2.28459 | 9.52013 | 1.0894 | -2.108041632
160 | 9.74677 | 0.991234 | 9.84743 | 1.73396 | 1.032752389
320 | 10.7297 | 5.11555 | 10.547 | 1.97692 | **-1.70275031**
640 | 11.7092 | 2.36565 | 11.7869 | 2.69377 | **0.6635807741**
- db bench on write with cost to cache in WriteBufferManager (just in case this PR's CRM refactoring accidentally slows down anything in WBM) : `fillseq` : **+0.54% in micros/op**
`./db_bench -benchmarks=fillseq -db=/dev/shm/testdb -disable_auto_compactions=1 -cost_write_buffer_to_cache=true -write_buffer_size=10000000000 | egrep 'fillseq'`
#-run | (pre-PR) avg micros/op | std micros/op | (post-PR) avg micros/op | std micros/op | change (%)
-- | -- | -- | -- | -- | --
10 | 6.15 | 0.260187 | 6.289 | 0.371192 | 2.260162602
20 | 7.28025 | 0.465402 | 7.37255 | 0.451256 | 1.267813605
40 | 7.06312 | 0.490654 | 7.13803 | 0.478676 | **1.060579461**
80 | 7.14035 | 0.972831 | 7.14196 | 0.92971 | **0.02254791432**
- filter bench: `bloom filter`: **-0.78% in ms/key**
- ` ./filter_bench -impl=2 -quick -reserve_table_builder_memory=true | grep 'Build avg'`
#-run | (pre-PR) avg ns/key | std ns/key | (post-PR) ns/key | std ns/key | change (%)
-- | -- | -- | -- | -- | --
10 | 26.4369 | 0.442182 | 26.3273 | 0.422919 | **-0.4145720565**
20 | 26.4451 | 0.592787 | 26.1419 | 0.62451 | **-1.1465262**
- Crash test `python3 tools/db_crashtest.py blackbox --reserve_table_reader_memory=1 --cache_size=1` killed as normal
Reviewed By: ajkr
Differential Revision: D35136549
Pulled By: hx235
fbshipit-source-id: 146978858d0f900f43f4eb09bfd3e83195e3be28
|
3 years ago |
.. |
adaptive
|
More refactoring ahead of footer & meta changes (#9240)
|
3 years ago |
block_based
|
Account memory of big memory users in BlockBasedTable in global memory limit (#9748)
|
3 years ago |
cuckoo
|
Add rate limiter priority to ReadOptions (#9424)
|
3 years ago |
plain
|
Replace GetUserKey with ExtractUserKey (#9664)
|
3 years ago |
block_fetcher.cc
|
Fix segfault in FilePrefetchBuffer with async_io enabled (#9777)
|
3 years ago |
block_fetcher.h
|
More refactoring ahead of footer & meta changes (#9240)
|
3 years ago |
block_fetcher_test.cc
|
Make MemoryAllocator into a Customizable class (#8980)
|
3 years ago |
cleanable_test.cc
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
format.cc
|
Add rate limiter priority to ReadOptions (#9424)
|
3 years ago |
format.h
|
Optimize & clean up footer code (#9280)
|
3 years ago |
get_context.cc
|
Support readahead during compaction for blob files (#9187)
|
3 years ago |
get_context.h
|
Cleanup includes in dbformat.h (#8930)
|
3 years ago |
internal_iterator.h
|
Reuse internal auto readhead_size at each Level (expect L0) for Iterations (#9056)
|
3 years ago |
iter_heap.h
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
iterator.cc
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
iterator_wrapper.h
|
Reuse internal auto readhead_size at each Level (expect L0) for Iterations (#9056)
|
3 years ago |
merger_test.cc
|
Cleanup multiple implementations of VectorIterator (#8901)
|
3 years ago |
merging_iterator.cc
|
MergingIterator: rearrange fields to reduce paddings (#9024)
|
3 years ago |
merging_iterator.h
|
Cleanup includes in dbformat.h (#8930)
|
3 years ago |
meta_blocks.cc
|
Add NewMetaDataIterator method (#8692)
|
3 years ago |
meta_blocks.h
|
More refactoring ahead of footer & meta changes (#9240)
|
3 years ago |
mock_table.cc
|
Add rate limiter priority to ReadOptions (#9424)
|
3 years ago |
mock_table.h
|
Fix some minor issues in the Customizable infrastructure (#8566)
|
3 years ago |
multiget_context.h
|
Fix major bug with MultiGet, DeleteRange, and memtable Bloom (#9453)
|
3 years ago |
persistent_cache_helper.cc
|
New stable, fixed-length cache keys (#9126)
|
3 years ago |
persistent_cache_helper.h
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
persistent_cache_options.h
|
New stable, fixed-length cache keys (#9126)
|
3 years ago |
scoped_arena_iterator.h
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
sst_file_dumper.cc
|
New backup meta schema, with file temperatures (#9660)
|
3 years ago |
sst_file_dumper.h
|
New backup meta schema, with file temperatures (#9660)
|
3 years ago |
sst_file_reader.cc
|
Fast path for detecting unchanged prefix_extractor (#9407)
|
3 years ago |
sst_file_reader_test.cc
|
Use the comparator from the sst file table properties in sst_dump_tool (#9491)
|
3 years ago |
sst_file_writer.cc
|
`compression_per_level` should be used for flush and changeable (#9658)
|
3 years ago |
sst_file_writer_collectors.h
|
Cleanup includes in dbformat.h (#8930)
|
3 years ago |
table_builder.h
|
Fast path for detecting unchanged prefix_extractor (#9407)
|
3 years ago |
table_factory.cc
|
Restore Regex support for ObjectLibrary::Register, rename new APIs to allow old one to be deprecated in the future (#9362)
|
3 years ago |
table_properties.cc
|
Account memory of big memory users in BlockBasedTable in global memory limit (#9748)
|
3 years ago |
table_properties_internal.h
|
Improve / clean up meta block code & integrity (#9163)
|
3 years ago |
table_reader.h
|
dedup ReadOptions in iterator hierarchy (#7210)
|
4 years ago |
table_reader_bench.cc
|
Fast path for detecting unchanged prefix_extractor (#9407)
|
3 years ago |
table_reader_caller.h
|
Fix and detect headers with missing dependencies (#8893)
|
3 years ago |
table_test.cc
|
Remove BlockBasedTableOptions.hash_index_allow_collision (#9454)
|
3 years ago |
two_level_iterator.cc
|
Clarify caching behavior for index and filter partitions (#9068)
|
3 years ago |
two_level_iterator.h
|
Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433)
|
5 years ago |
unique_id.cc
|
Experimental support for SST unique IDs (#8990)
|
3 years ago |
unique_id_impl.h
|
Experimental support for SST unique IDs (#8990)
|
3 years ago |