Console View
|
Categories: connectors experimental galera main |
|
| connectors | experimental | galera | main | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
|
|
|
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
bsrikanth-mariadb
srikanth.bondalapati@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40837 Fix crash recording opt context with NULL character_set_results Unlike character_set_client or collation_connection, character_set_results can be set to NULL. When recording context for a query in opt_context_store_replay.cc, character_set_results->cs_name was accessed without first checking whether character_set_results itself was NULL, causing a crash. Fix Optimizer_context_recorder::dump_sql_script() to check character_set_results for NULL before accessing cs_name, writing "SET character_set_results=NULL;" to the replay script in that case so replay reproduces the recorded session instead of silently defaulting to the script's SET NAMES charset. Added a test for the same, in opt_context_store_stats.test, that also checks the recorded script contains the NULL SET statement. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| s | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alexander Barkov
bar@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41246 "Illegal mix of collations" on the mysql.user view In progress |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alexey (Holyfoot) Botchkov
holyfoot@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-39923 HANDLER read crashes restore_position() for InnoDB partitioned table. m_top_entry has to be reset in ha_partition::index_init(). A stale value left over from an earlier ordered scan on this handler breaks the execution. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41322 Convert KEY_NOT_FOUND to EOF in pq-less partition index scans during index_next/index_prev calls ha_partition::handle_unordered_scan_next_partition is called in a variety of accesses, including index_read, index_prev, and index_next. ha_partition::handle_unordered_next and ha_partition::handle_unordered_prev are called from index_next and index_prev accesses. They check for signs of end of scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever possible, to signal the end of scan. The error HA_ERR_KEY_NOT_FOUND means the requested key is not found. It should not mean the end of scan, when for example ha_partition::handle_unordered_scan_next_partition is called from index_read, because a subsequent index_next / index_prev call would then incorrectly return immediately from ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND is retained and returned in ha_partition::handle_unordered_scan_next_partition. But if the call is from index_next / index_prev, HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-41322 Convert HA_ERR_KEY_NOT_FOUND to HA_ERR_END_OF_FILE in pq-less partition index scan | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
DerZc
34330257+DerZc@users.noreply.github.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40820 Wrong results: SELECT DISTINCT / GROUP BY returns duplicate rows on a RANGE-partitioned table when served by a covering index range scan SELECT DISTINCT or GROUP BY can return duplicate values from a RANGE- partitioned table when a covering index range scan incorrectly bypasses the merge of partition scans. ha_partition::can_skip_merging_scans() checks only the current multi- range prefix. Later ranges can have different prefix values, so the partition outputs do not have the ordering required to skip the priority-queue merge. Check every multi-range entry before bypassing the partition-scan merge. Require both endpoints to bind the complete unordered prefix and to agree on its bytes. Require that prefix to be the same across all ranges; otherwise keep the normal merge. The regression uses two date prefixes across several partitions and checks that SELECT DISTINCT returns each date exactly once. Bug report: https://jira.mariadb.org/browse/MDEV-40820 |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41322 Convert KEY_NOT_FOUND to EOF in "unordered" partition index scans during index_next[_same]/index_prev calls ha_partition::handle_unordered_scan_next_partition is called in a variety of accesses, including index_read, index_prev, and index_next. ha_partition::handle_unordered_next and ha_partition::handle_unordered_prev are called from index_next and index_prev accesses. They check for signs of end of scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever possible, to signal the end of scan. The error HA_ERR_KEY_NOT_FOUND means the requested key is not found. It should not mean the end of scan, when for example ha_partition::handle_unordered_scan_next_partition is called from index_read, because a subsequent index_next[_same] / index_prev call would then incorrectly return immediately from ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND is retained and returned in ha_partition::handle_unordered_scan_next_partition. But if the call is from index_next[_same] / index_prev, HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-41322 Convert HA_ERR_KEY_NOT_FOUND to HA_ERR_END_OF_FILE in pq-less partition index scan | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
mariadb-PranavTiwari
pranav.tiwari@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Added mysql_upgrade | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-38347 Debugging functions Extend current dbug_print* functions in such a way as to facilitate easy high level debugging. Introduce a single overloaded print function defined in dbp.h, DBUG_PRINT_FUNCTION, settable to your preference. Allow a shell environment variable DBUG_PRINT_FUNCTION to override this definition |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41157 Follow-up test fix: strip version-dependent charset result Follow-up to 6c391b92055. SHOW CREATE DATABASE always appends "/*!40100 DEFAULT CHARACTER SET ... */", but the default charset differs by version (latin1 on 11.4, utf8mb4 on 11.8+). Strip it instead of matching a specific value, since it is not the subject for the fix. Also switch to evalp: same execution, but logs "$var" instead of the substituted value, so the replace_regex/disable_query_log hacks go away. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41322 Convert KEY_NOT_FOUND to EOF in pq-less partition index scans during index_next/index_prev calls ha_partition::handle_unordered_scan_next_partition is called in a variety of accesses, including index_read, index_prev, and index_next. ha_partition::handle_unordered_next and ha_partition::handle_unordered_prev are called from index_next and index_prev accesses. They check for signs of end of scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever possible, to signal the end of scan. The error HA_ERR_KEY_NOT_FOUND means the requested key is not found. It should not mean the end of scan, when for example ha_partition::handle_unordered_scan_next_partition is called from index_read, because a subsequent index_next / index_prev call would then incorrectly return immediately from ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND is retained and returned in ha_partition::handle_unordered_scan_next_partition. But if the call is from index_next / index_prev, HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Vladislav Vaintroub
vvaintroub@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3 OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1 and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can stop matching after switching to another. Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts into falling back to an escape-aware compare, applied to whichever side the currently-linked library's own escaping affects, at the cost of reopening the ambiguity a crafted certificate could exploit to impersonate another identity. Adds regression tests against a real certificate with an ambiguous CN Assisted-by: Claude Sonnet 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Jan Lindström
jan.lindstrom@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40622 : galera.tmp_space_usage fails: Failed to start mysqld.2 Max_tmp_space_used and Tmp_space_used are binlog cache byte counts that vary by platform. The test compared them against fixed numbers, so a differing byte count failed the test. The test now checks that both values stay within max_tmp_session_space_usage, and that tmp_space_used resets to 0 after change_user. Test only. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Vladislav Vaintroub
vvaintroub@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3 OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1 and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can stop matching after switching to another. Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts into falling back to an escape-aware compare, applied to whichever side the currently-linked library's own escaping affects, at the cost of reopening the ambiguity a crafted certificate could exploit to impersonate another identity. Adds regression tests against a real certificate with an ambiguous CN Assisted-by: Claude Sonnet 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41050 versioned DELETE via row_end index leaves a row undeleted 1. Aria/MyISAM engines A system-versioned DELETE on a MyISAM or Aria table indexed by row_end could leave the last current row undeleted. The delete scans the current rows via an equality search on row_end = MAX, which is driven by mi_rnext_same/maria_rnext_same. That function keeps the search's reference key in lastkey2 and uses the HA_STATE_RNEXT_SAME flag to remember it has already stored it. This reference-key mechanism is exactly what lets the scan modify rows it is walking without skipping them. Deleting a versioned row is an in-place update of row_end, and the update reuses lastkey2 as scratch space for the changed key, so it clears HA_STATE_RNEXT_SAME to request that rnext_same re-store its reference on the next call. However, TABLE::delete_row wraps the update in HA_EXTRA_REMEMBER_POS/HA_EXTRA_RESTORE_POS, and RESTORE_POS restored the whole saved info->update word, resurrecting the HA_STATE_RNEXT_SAME bit that the update had just cleared. As a result rnext_same skipped rebuilding its reference key and compared subsequent keys against the now-overwritten lastkey2, hitting a spurious end-of-file and terminating the scan one row early, defeating the engine's own protection against a Halloween-style skip. Fixed by preserving the current HA_STATE_RNEXT_SAME bit across RESTORE_POS instead of restoring the stale saved value. See also the HEAP fix below: same root cause, different per-engine mechanism. 2. HEAP engine The row loss also reproduces on the MEMORY (HEAP) engine. A system- versioned DELETE scans the current rows on the row_end index and turns each delete into an in-place update of row_end, so it modifies the very index it is walking. When the changed key is the scanned one (info->lastinx), hp_delete_key() repositions the cursor but heap_update() leaves info->update untouched, so HA_STATE_NEXT_FOUND from the preceding heap_rnext() stays set. The next heap_rnext() then sees current_ptr == 0 with that bit and takes the "!current_ptr && HA_STATE_NEXT_FOUND" guard as a false end-of-file, stopping one row early. Fixed by clearing HA_STATE_NEXT_FOUND when the scanned index key changed. HA_STATE_AKTIV is kept (unlike heap_delete): the row is updated, not removed, so a following op must not fail test_active(). The bit is only set after a heap_rnext(), so a plain single-row UPDATE never reaches this. See also the Aria/MyISAM fix above: same root cause, different per-engine mechanism. 3. Why the fix differs per engine, and InnoDB Aria and HEAP both trace their handler design back to MyISAM, hence the same root cause (stale scan bookkeeping after an in-place key change) in all three, fixed at each engine's own bookkeeping spot. InnoDB needs no fix: its persistent cursor survives concurrent index modification by design, already covered by this same test under the timestamp combination (default-storage-engine=innodb), which passes unmodified. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Vladislav Vaintroub
vvaintroub@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3 OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1 and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can stop matching after switching to another. Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts into falling back to an escape-aware compare, applied to whichever side the currently-linked library's own escaping affects, at the cost of reopening the ambiguity a crafted certificate could exploit to impersonate another identity. Adds regression tests against a real certificate with an ambiguous CN Assisted-by: Claude Sonnet 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41322 Convert KEY_NOT_FOUND to EOF in "unordered" partition index scans during index_next[_same]/index_prev calls ha_partition::handle_unordered_scan_next_partition is called in a variety of accesses, including index_read, index_prev, and index_next. ha_partition::handle_unordered_next and ha_partition::handle_unordered_prev are called from index_next and index_prev accesses. They check for signs of end of scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever possible, to signal the end of scan. The error HA_ERR_KEY_NOT_FOUND means the requested key is not found. It should not mean the end of scan, when for example ha_partition::handle_unordered_scan_next_partition is called from index_read, because a subsequent index_next[_same] / index_prev call would then incorrectly return immediately from ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND is retained and returned in ha_partition::handle_unordered_scan_next_partition. But if the call is from index_next[_same] / index_prev, HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in ha_partition::handle_unordered_next / ha_partition::handle_unordered_prev. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41157 Follow-up test fix: strip version-dependent charset result Follow-up to 6c391b92055. SHOW CREATE DATABASE always appends "/*!40100 DEFAULT CHARACTER SET ... */", but the default charset differs by version (latin1 on 11.4, utf8mb4 on 11.8+). Strip it instead of matching a specific value, since it is not the subject for the fix. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Vladislav Vaintroub
vvaintroub@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3 OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1 and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can stop matching after switching to another. Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts into falling back to an escape-aware compare, applied to whichever side the currently-linked library's own escaping affects, at the cost of reopening the ambiguity a crafted certificate could exploit to impersonate another identity. Adds regression tests against a real certificate with an ambiguous CN Assisted-by: Claude Sonnet 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Thirunarayanan Balathandayuthapani
thiru@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41207 Acquire metadata locks for recovered transaction Problem: ======= A transaction being rolled back during recovery holds LOCK_IX on the table(not the metadata locks), and the rollback thread holds a reference on it. An online ALTER TABLE on that table falls back to acquiring LOCK_S, which conflicts and fails. prepare_inplace_alter_table_dict() asserted that the reference count is 1 before checking whether the table lock was acquired, so the reference still held by the rollback thread makes the assertion fail. Solution: ======== recovery_mdl: New class that holds the metadata locks which trx_resurrect_table_locks() acquires for recovered transactions, together with the background connection that owns them. One MDL_ticket is acquired per table and shared by every recovered transaction that modified the table, with a reference count, so that the lock is released once the last of those transactions has been completed. trx_sys_t::recovery: Pointer to recovery_mdl, wrapped in std::atomic, next to rw_trx_hash, which trx_t::free() accesses anyway. Null pointer means that no recovered transaction holds metadata locks, which is the normal state once recovery is over. trx_lists_init_at_db_start(): Create recovery_mdl and publish it in trx_sys.recovery before any metadata lock can be acquired, and delete it again if no recovered transaction acquired any. recovery_mdl::release(): Drop the references of a transaction, releasing each metadata lock whose last reference is dropped. This is the only place that releases them, so that a transaction which was rolled back by trx_rollback_recovered(false) will not keep its locks until shutdown. trx_t::free(): Release the metadata locks of a recovered transaction, once it has been completed. A single relaxed load of trx_sys.recovery is all that an ordinary transaction pays. trx_recovery_release(): Re-read trx_sys.recovery while holding trx_recovery_mutex, because the object may have been deleted after trx_t::free() read the pointer. trx_recovery_prune_low(): Delete trx_sys.recovery once no recovered transaction holds metadata locks any longer, so that neither the locks nor the connection will outlive the recovered transactions. This is skipped when the caller has a current_thd, such as a user connection that is completing a transaction which was recovered in the XA PREPARED state, and during shutdown, because destroy_background_thd() would clear the current_thd of the caller. trx_recovery_shutdown(): Destroy trx_sys.recovery, releasing anything that is left. It is invoked by innodb_shutdown() only, after trx_sys.close() has freed any recovered transaction that was left. prepare_inplace_alter_table_dict(): Remove the acquisition of an InnoDB table lock for online ADD INDEX when a recovered transaction holds a lock on the table. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Rex Johnston
rex.johnston@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41179 MAX_KEY_PARTS ranges in select causing SEGV When a select containing 32 ranges is made on a table containing a compound key with 32 parts, the range optimizer can run off the end of a stack variable, invalidly overwriting subsequent stack variables. In the struct st_sel_arg_range_seq, we have an array RANGE_SEQ_ENTRY stack[MAX_REF_PARTS]; MAX_REF_PARTS is 32. check_quick_select / sel_arg_range_seq_init initialises stack[0] as NOT a key part / sel_arg_range_seq_next iterates through the key parts, adding key part n to stack[n+1] key part #32 gets referenced by step_down_to(), setting seq->i off the end of the array. The above change has exposed an issue with key length calculation on MS Windows. The field is a ulong and with all 32 key parts wanted, we calculate this as (1UL << 32) - 1. The first part of the calculation overflows the field. The result is undefined in C++. Fix: change key_part_map to 64 bits long, so it is the same on Linux and MS Windows. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
bsrikanth-mariadb
srikanth.bondalapati@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-36986: Support tracing array of primitive types Json_writer had two separate code paths: add_unquoted_str() for numbers/bool/null, and add_escaped_str() which added the surrounding quotes itself for strings. Single_line_formatting_helper, which buffers consecutive array/object elements to decide if they fit on one line, assumed only strings could ever be buffered and always wrapped the flushed values in quotes. As a result, arrays of numbers (e.g. "depends_on_map_bits", "rec_per_key") were incorrectly rendered with their elements quoted as strings. Unify both paths into add_escaped_quoted_str(): the caller now hands over bytes that are already in their final on-the-wire form. String escaping (json_escape_to_string) writes its own surrounding quotes, while numbers/bool/null are passed through unquoted, so the one-line helper just concatenates the buffered payloads on flush instead of adding quotes itself. A DBUG_ASSERT in add_escaped_quoted_str() now checks that every payload is already a quoted string or a bare number/bool/null token, so a caller that violates the contract trips an assertion in debug builds instead of silently producing invalid JSON. Also: - Fix Json_writer_array::add(ulonglong)/(size_t), which went through add_ll() with a cast to longlong and corrupted large unsigned values (e.g. ULLONG_MAX); route them through add_ull() instead. - Fix mysql-test/include/opt_context_schema.inc: "subquery_runs" was nested inside the preceding object instead of being a sibling member, and "rec_per_key" items are now declared as "number" to match the corrected output. - Update recorded .result files for opt_trace, opt_context_*, and subselect_mat_analyze_json to reflect numbers/booleans no longer being quoted inside JSON arrays. - Extend unittest/sql/my_json_writer-t.cc with coverage for arrays of primitives: plain integers, mixed types, sizes, multi-line arrays, values flushed before a nested object, strings that were already escaped by the one-line buffer, and invalid utf8mb4 input through add_str(). - Rewrite the "multi-line array of integers" test to actually overflow the one-line buffer (9 seven-digit numbers instead of 7), since the previous version fit on one line and never exercised the element-per-line flush path it claimed to test. - Drop the now-redundant single-argument add_escaped_quoted_str() overload; all call sites already know their length. - Update json_escape_to_string()'s doc comments (my_json_writer.h, sql_json_lib.h) to state that it quotes its output, not just escapes it. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Dave Gosselin
dave.gosselin@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-31180: MyISAMMRG Crash on UPDATE of an updateable VIEW Attach the children of a MERGE table once per statement, and keep the value of pos_in_table_list for a MERGE table on subsequent executions of a prepared statement. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alessandro Vetere
iminelink@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-39792 InnoDB: ALTER TABLE FORCE triggers assertion "s" in buf_page_get_gen() When rebuilding a table from ROW_FORMAT=COMPACT or DYNAMIC into ROW_FORMAT=REDUNDANT, row_merge_buf_add() fetches the full value of an externally stored (off-page) CHAR column in a multi-byte character set and pads it to REDUNDANT's fixed local width via row_merge_buf_redundant_convert(). That helper already dereferences the BLOB and calls dfield_set_data(), which clears the field's "externally stored" flag, since the value is now held in full locally. The "flag externally stored fields" step further down in row_merge_buf_add() did not know this had happened. It still consulted the row_ext_t cache built from the original (pre-conversion) record and, for a column that is not part of the clustered index's unique key, called dfield_set_ext() again on the very field that had just been converted, without restoring its data pointer to a valid 20-byte external reference. row_merge_copy_blobs() would then read the tail of the padded, space-filled buffer as if it were a BTR_EXTERN_FIELD_REF, deriving a garbage tablespace id and crashing buf_page_get_gen()'s fil_space_get() assertion when the alter tried to build the new clustered index. Skip the re-flagging step for a field whose "externally stored" flag is no longer set. row_build() flags every off-page column, and the row_ext_t cache only holds a subset of those columns, so a field that is not flagged is either a converted one (already fully local) or one that the cache does not hold. With the field no longer re-flagged, the rebuild completes, and the rebuilt table passes CHECK TABLE with the full column value. The MDEV-31025 case in innodb.default_row_format_alter failed on innodb_page_size=4k and 8k: its ROW_FORMAT=REDUNDANT table has eight utf32 CHAR(255) columns, which CREATE TABLE rejects with ER_TOO_BIG_ROWSIZE on those page sizes. Derive the number of columns from the page size, so that the record still exceeds the maximum local record size and the fixed-length column c is stored externally. The whole test now passes on every page size. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
bsrikanth-mariadb
srikanth.bondalapati@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41243, MDEV-41273: update, delete, select prepared statements assert with fedx_pushdown PROTECT_STATEMENT_MEMROOT freezes a statement's/routine's permanent arena read-only once it believes every once-per-statement optimization (SELECT_LEX::first_cond_optimization, run inside JOIN::optimize_inner()) has had its chance to execute. Pushdown breaks that assumption: when a SELECT/UPDATE/DELETE/UNION's execution is handed off to a storage engine, JOIN::optimize_inner() (and everything it would have done) is skipped entirely. If that h, the arena gets frozen early; a later execution that is no longer pushed down then tries to run that defest time against an already-frozen arena and hits: alloc_root: Assertion `(mem_root->flags & 4) == 0' failed Fix: track, per LEX, whether a pushed-down execution left a once-per-statement optimiza (LEX::pushdown_skipped_first_execution_optimization), and defer freezing until an executione that hands execution to a storage engine without going through JOIN::optimize_inner() has ble; every place that would otherwise freeze an arena has to check it first. - sql_class.h: JOIN::optimize() sets the flag on the currently executing LEX when a SELEion is still pending at the point pushdown takes over. - Living on LEX rather thanement and each stored-routine instruction (each with its own persistent LEX, reused across executions until rently, with no save/restore needed across nested CALL/EXECUTE boundaries. - sql_prepare.cc: Prepared_ers freezing its own mem_root while the flag is set. - sp_head.cc/sp_instr.h: spction loop defers freezing main_mem_root the same way, via two new sp_instr virtual methods overriddestr_copen (OPEN cursor_name) has no LEX of its own -- it runs the cursor's SELECT using the corresponding Duction's LEX -- so it overrides them to check there instead of defaulting to "nothing pending". - sp_instr.cc: sp_lex_keeper::validate_lex_and_exec_core() defers freezing a stored-routinereparse mem_root the same way, carrying the deferral across calls via a new m_mem_root_freeze_pending first post-reparse execution happens to be pushed down. - sql_union.cc: a UNION puss optimize()/exec_inner() (and thus JOIN::optimize()) entirely, so st_select_lex_unit gains ptimization(), checking every member SELECT, fake_select_lex (the UNION's own ORDER BY/LIMIT/DISTINCT wy unit nested anywhere inside either one, since a pushdown hand-off gives away that whole subtree, not just t - sql_select.cc/sql_derived.cc: the same check is needed wherever else a unit's execution can bereaching JOIN::optimize() -- EXPLAIN of a pushed unit (mysql_explain_union()) and a UNION used as a derl()'s pushdown_derived branch). Test: federated.federatedx_pushdown_upd_del, covering top-level prepared UPDATE/DELETE/SELE UNION, a UNION used as a derived table (including a UNION nested inside a subquery inside that derived table), stored routine (CALL, nested CALL, EXECUTE of a PS from inside a routine, a static cursor, and a metada |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Mohammad Tafzeel Shams
tafzeel.shams@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-37467: InnoDB Instant ALTER TABLE is not crash safe The hidden metadata record of instant ALTER TABLE was not written crash-safely, and recovery could fail to roll it back. These are independent problems. First, the metadata record may include externally stored BLOB metadata. The existing BLOB storage path in btr_store_big_rec_extern_fields() writes the clustered index record first, with zero BLOB pointers, and only fills in the BLOB pointers afterwards. If the server is killed after the mini-transaction that wrote the (incomplete) metadata record was durably committed, but before the BLOB pointers were written, the table could become inaccessible on recovery. Make metadata BLOB storage crash-safe by writing the BLOB pages and computing their pointers before the metadata record itself is inserted or updated, so that the record is always written with complete BLOB pointers. If the server is killed before the metadata record is written, the already-written BLOB pages are merely orphaned, which is safe. Second, trx_undo_report_row_operation() writes the undo log record in a mini-transaction of its own, which is committed before the mini-transaction that writes the metadata record. Because innobase_instant_try() had already updated SYS_COLUMNS and SYS_TABLES in earlier mini-transactions, a kill in between left a durable undo log record for the table while the metadata record was unchanged. On recovery, trx_resurrect_table_locks() would then load the table definition before the incomplete transaction was rolled back. The data dictionary described the table as it would be after the operation, while the metadata record still described it as it was before, and btr_cur_instant_init() failed on that disagreement. Write the undo log record of the metadata record in the same mini-transaction that inserts or updates the record, so that the two cannot be separated by a crash: until that mini-transaction is committed, neither of them is durable. An undo log record is never split between pages. If the DEFAULT values of the columns being added are large enough that the undo log record for updating the metadata record would not fit on one page, innobase_instant_try() would fail. Determine this before the operation starts, so that it can be performed by another algorithm instead. Third, the table definition that recovery loads need not correspond to the metadata record. dict_load_table_one() reads the committed version of the SYS_TABLES record, and escalates to READ UNCOMMITTED only when it finds a SYS_COLUMNS record that was written by a transaction that is still active. The number of SYS_COLUMNS records that dict_load_columns() reads is derived from SYS_TABLES.N_COLS, which was read from the committed version. The record of a column that the operation appended is located after that many records, so it is never read and the operation goes unnoticed. Only an instant ALTER TABLE that merely appends columns can escape this way: ADD COLUMN ... FIRST, DROP COLUMN and column reordering rewrite the SYS_COLUMNS records of already existing columns. Detect this on the SYS_TABLES record itself, which is located by table name and therefore does not depend on N_COLS. Every instant ALTER TABLE that changes the columns updates that record, because innobase_instant_try() invokes innodb_update_cols(). Fourth, the rollback writes a metadata record that comprises fewer fields than the table definition describes, because btr_cur_trim_alter_metadata() shortens it to the number of fields that it comprised before the operation. That number determines the size of the null flag bitmap, and hence the position of the array of field lengths. rec_init_offsets_comp_ordinary() derives it from the record, while the two functions that write the record derived it from the table definition and asserted that the two agree. - btr_store_big_rec_metadata(): New function to store the off-page columns of a metadata record ahead of time. Each BLOB page is allocated and linked in its own mini-transaction, and the resulting BLOB pointers are written directly into the (heap-resident) index entry. On failure, it frees any pages it already allocated and resets the pointers to zero. - btr_free_big_rec_metadata(): New helper to free the BLOB pages written by btr_store_big_rec_metadata() and reset the entry's BLOB pointers to zero, used both on failure inside that function and by its callers when the metadata record ends up not being written. - row_ins_clust_index_entry_low(): For a metadata entry that needs external storage, convert it to a big record and call btr_store_big_rec_metadata() (with log_free_check() allowed, since no latches are held yet) before inserting the record. On failure, free the metadata BLOBs and convert the entry back. - btr_cur_pessimistic_update(): When updating a metadata record that requires external storage, call btr_store_big_rec_metadata() (without log_free_check(), since index and page latches are held) before modifying the record, and free the temporary big_rec vector via btr_free_big_rec_metadata() or dtuple_big_rec_free() on the various failure/success paths. - btr_cur_optimistic_insert(): Remove the special-cased jump to convert_big_rec for metadata entries, since their BLOBs are now always stored ahead of time by the caller; assert that a metadata entry never needs external storage at this point. - innobase_instant_try(): Since btr_cur_pessimistic_update() now stores metadata BLOBs before updating the record, big_rec is always NULL here; assert this instead of calling btr_store_big_rec_extern_fields(). - trx_undo_report_row_operation(): New parameter caller_mtr. If it is specified, the undo log record is written in that mini-transaction, which is never committed or restarted here. An undo log page is added within the same mini-transaction if the record does not fit on the current one. A temporary table never uses the caller's mini-transaction, because that would require changing its logging mode. All other callers pass NULL and are unaffected. Encapsulate the parameters that describe the row change (clust_entry, update, cmpl_info, rec, offsets) in the new type trx_undo_row_op, which of the fields are set depends on the operation, which the type documents. - btr_cur_ins_lock_and_undo(), btr_cur_upd_lock_and_undo(): For an instant ALTER TABLE metadata record, pass the mini-transaction that is going to insert or modify the record. - trx_undo_max_rec_size(): New function to determine the maximum size of an undo log record, that is, the space available on an empty undo log page. - ha_innobase::check_if_supported_inplace_alter(): Refuse ALGORITHM=INSTANT if the metadata record already exists and the undo log record for updating it would exceed trx_undo_max_rec_size(). trx_undo_page_report_modify() stores the DEFAULT value of each column that is being added in that record in full, inline. No such limit applies when the metadata record is being inserted, because trx_undo_page_report_insert() writes TRX_UNDO_INSERT_METADATA and no field data. - dict_load_table_one(): If the SYS_TABLES record was written by a transaction that is still active, load the table definition as READ UNCOMMITTED. A delete-marked record is excluded, because SYS_TABLES.NAME is the clustered index key: RENAME TABLE delete-marks the record of the old name, and the definition that corresponds to that name is the one that precedes the rename. - dict_sys_tables_rec_read(): Report whether the current version of the record was written by a transaction that has not been committed. The function already determines this in order to decide whether to read an older version of the record, and used to discard the answer. Store the fields that are read in the new type dict_sys_tables_rec, instead of in separate output parameters. A caller that does not want a delete-marked record to be reported as not found used to indicate that by specifying trx_id as nullptr; that is now the parameter skip_deleted. - dict_load_table_low(): New parameter uncommitted_rec, which is passed on to dict_sys_tables_rec_read(). - rec_get_converted_size_comp_prefix_low(), rec_convert_dtuple_to_rec_comp(): For a record that includes a metadata BLOB, determine the number of nullable fields from the tuple, by way of dict_index_t::get_n_nullable(), and not from dict_index_t::n_nullable. This is what rec_init_offsets_comp_ordinary() does, and it is equivalent for a tuple that comprises all fields of the index. Relax the assertions that required the tuple to comprise all of them. - Added test in innodb.instant_alter and innodb.instant_alter_crash to test normal working of INSTANT ALTER, crash safety and full table. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||