Console View
|
Categories: connectors experimental galera main |
|
| connectors | experimental | galera | main | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
|
|
|
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
bsrikanth-mariadb
srikanth.bondalapati@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40553 Print GIS ranges in optimizer trace and context Ranges built over GIS (geometry) columns could not be printed in the optimizer trace or the recorded optimizer context: Field_geom printed every key value as the placeholder "unprintable_geometry_value", regardless of whether the index stored the column's raw value (or a prefix of it) or, for a SPATIAL index, its MBR (Minimum Bounding Rectangle). Field::print_key_part_value() now takes an image_type argument (see Field::image_type()) that says which of the two the key holds. For a SPATIAL index (image_type itMBR), Field_geom::print_key_part_value() decodes the four doubles the key stores and prints them as a WKT POLYGON. For every other index (image_type itRAW), the key holds the value or a prefix of it, and it now prints in binary form, the same way Field_blob does, instead of the placeholder. print_mbr_range_operator() prints the spatial relation a GEOM range carries (MBRWITHIN, MBRCONTAINS, MBRINTERSECTS, MBRDISJOINT, MBREQUALS), inverted where needed so the indexed column reads on the left. print_range() and print_key_value() thread the new image_type argument through to Field_geom::print_key_part_value(). Added new tests in opt_trace, opt_context_store_stats, opt_context_replay_basic |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Allow a covering index scan for a SELECT ... FOR UPDATE under a full-scan table lock (TODO: improve commit message) MDEV-24813 introduced a switch innodb_table_lock_on_full_scan that places a table lock on an innodb table for some unconditional SELECT statements. With a full table lock, and when doing a covering scan, there is no need to acquire clustered locks. This patch does so for X-locks. TODO: add code comments and tests Co-Authored-By: Claude Opus 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Oleksandr Byelkin
sanja@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40303 SIGSEGV in Item_field::type_handler() on PS re-execution On PS re-execution a derived column is resolved via item->cached_table. If fixing its scalar UNION subquery fails, the unit's prepare() error path calls cleanup() but leaves 'prepared' set. find_field_in_view() returned 0 ("not found"), so find_field_in_tables() fell through to the generic lookup, re-fixed the same subquery, prepare() returned early and set_row() hit NULL fields. Fix: return view_field_error when a view/derived field cannot be fixed and stop the lookup on it. Co-Authored-By: Claude Opus 5.5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Oleksandr Byelkin
sanja@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41099 HANDLER READ missing column privilege check HANDLER OPEN accepted column-level SELECT grants as table access, but HANDLER READ returns all columns and never checked column privileges. Reopen of a handler table (after FLUSH/ALTER) did no checks at all. Require SELECT on every column (as for SELECT *) when the user has no table-level SELECT, and redo the privilege checks on reopen. Co-Authored-By: Claude Opus 5.5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Marko Mäkelä
marko.makela@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
fixup! 33b746888e75892654df7ffa3a696512656a631d Calculate the correct file offset when zeroing the unused tail. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Georgi (Joro) Kodinov
joro@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| more doxygen formatting added. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Andrei Elkin
andrei.elkin@pp.inet.fi |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41214 PURGE BINARY LOGS does not delete non-active binlog files PURGE BINARY LOGS could end with a warning instead of deleting files, as the test below shows when run on the base of this commit. The warning names a wrong reason. It follows an automatic purge that was refused by the slave guard, slave_connections_needed_for_purge: the automatic purge does not delete a binlog while it might still be needed by a slave, not necessarily a connected one, that is while fewer than the variable's number of slaves are connected. Since MDEV-34504 the manual PURGE is meant to ignore the guard. Yet it still protects a binlog that a slave is reading. Technically this was possible because the automatic and the manual purge share a code path, can_purge_log(), which cached the refusal and replayed it, under the reason of the first arm below, to any later purge: if (is_active(name) || (!is_relay_log && waiting_for_slave_to_change_binlog && purge_sending_new_binlog_file == sending_new_binlog_file && !strcmp(name, purge_binlog_name))) reason= "it is the current active binlog"; So "active" had to be read as "active or busy". The fix removes this duality too: the two arms are now separate and have their own reasons. The main issue is fixed by narrowing the busy arm with an additional `!interactive &&` conjunct, so that a manual PURGE never consults the cached refusal. Running the test below on an unfixed 11.4 server (source tree at commit 42038e1145d) fails with "Result length mismatch" and, in part 2, this diff against the recorded result: PURGE BINARY LOGS TO 'master-bin.000006'; +Warnings: +Note 1375 Binary log 'master-bin.000003' is not purged because it is the current active binlog SHOW WARNINGS; Level Code Message +Note 1375 Binary log 'master-bin.000003' is not purged because it is the current active binlog show binary logs; Log_name File_size +master-bin.000003 # +master-bin.000004 # +master-bin.000005 # master-bin.000006 # master-bin.000007 # master-bin.000008 # Test binlog.binlog_purge_stale_refusal_cache: the test sets slave_connections_needed_for_purge=1, binlog_expire_logs_seconds=0 and max_binlog_total_size=0 first and restores them at the end. It starts with RESET MASTER so that the binlog names, 000003 and so on, do not depend on the tests run before it in the same mtr environment. Part 1 is a control with nothing cached; part 2 makes an automatic purge refuse 000003 (the files must be older than binlog_expire_logs_seconds, hence --real_sleep 2 before the rotation) and then runs PURGE BINARY LOGS TO 'master-bin.000006', which must remove 000003..000005. As a workaround, a PURGE could work if the following settings are made before it, in this order: SET GLOBAL binlog_expire_logs_seconds=0, slave_connections_needed_for_purge=0; The second setting invalidates the cached refusal. The first one keeps the purge that this setting triggers from refusing and caching again. Note that both settings change the server's policy, not only this one PURGE: the time-based expiry and the slave guard are off until they are restored. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Monty
monty@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-25292 Atomic CREATE OR REPLACE TABLE Atomic CREATE OR REPLACE allows to keep an old table intact if the command fails or during the crash. That is done by renaming the original table to temporary name, as a backup and restoring it if the CREATE fails. When the command is complete and logged the backup table is deleted. Atomic replace algorithm Two DDL chains are used for CREATE OR REPLACE: ddl_log_state_create (C) and ddl_log_state_rm (D). 1. (C) Log rename of ORIG to TMP table (Rename TMP to original). 2. Rename orignal to TMP. 3. (C) Log CREATE_TABLE_ACTION of ORIG (drops ORIG); 4. Do everything with ORIG (like insert data) 5. (D) Log drop of TMP 6. Write query to binlog (this marks (C) to be closed in case of failure) 7. Execute drop of TMP through (D) 8. Close (C) and (D) If there is a failure before 6) we revert the changes in (C) Chain (D) is only executed if 6) succeded (C is closed on crash recovery). Foreign key errors will be found at the 1) stage. Additional notes - CREATE TABLE without REPLACE and temporary tables is not affected by this commit. set @@drop_before_create_or_replace=1 can be used to get old behaviour where existing tables are dropped in CREATE OR REPLACE. - CREATE TABLE is reverted if binlogging the query fails. - Engines having HTON_EXPENSIVE_RENAME flag set are not affected by this commit. Conflicting tables marked with this flag will be deleted with CREATE OR REPLACE. - Replication execution is not affected by this commit. - Replication will first drop the conflicting table and then creating the new one. - CREATE TABLE .. SELECT XID usage is fixed and now there is no need to log DROP TABLE via DDL_CREATE_TABLE_PHASE_LOG (see comments in do_postlock()). XID is now correctly updated so it disables DDL_LOG_DROP_TABLE_ACTION. Note that binary log is flushed at the final stage when the table is ready. So if we have XID in the binary log we don't need to drop the table. - Three variations of CREATE OR REPLACE handled: 1. CREATE OR REPLACE TABLE t1 (..); 2. CREATE OR REPLACE TABLE t1 LIKE t2; 3. CREATE OR REPLACE TABLE t1 SELECT ..; - Test case uses 6 combinations for engines (aria, aria_notrans, myisam, ib, lock_tables, expensive_rename) and 2 combinations for binlog types (row, stmt). Combinations help to check differences between the results. Error failures are tested for the above three variations. - expensive_rename tests CREATE OR REPLACE without atomic replace. The effect should be the same as with the old behaviour before this commit. - Triggers mechanism is unaffected by this change. This is tested in create_replace.test. - LOCK TABLES is affected. Lock restoration must be done after new table is created or TMP is renamed back to ORIG - Moved ddl_log_complete() from send_eof() to finalize_ddl(). This checkpoint was not executed before for normal CREATE TABLE but is executed now. - CREATE TABLE will now rollback also if writing to the binary logging failed. See rpl_gtid_strict.test backup ddl log changes - In case of a successfull CREATE OR REPLACE we only log the CREATE event, not the DROP TABLE event of the old table. ddl_log.cc changes ddl_log_execute_action() now properly return error conditions. ddl_log_disable_entry() added to allow one to disable one entry. The entry on disk is still reserved until ddl_log_complete() is executed. On XID usage Like with all other atomic DDL operations XID is used to avoid inconsistency between master and slave in the case of a crash after binary log is written and before ddl_log_state_create is closed. On recovery XIDs are taken from binary log and corresponding DDL log events get disabled. That is done by ddl_log_close_binlogged_events(). On linking two chains together Chains are executed in the ascending order of entry_pos of execute entries. But entry_pos assignment order is undefined: it may assign bigger number for the first chain and then smaller number for the second chain. So the execution order in that case will be reverse: second chain will be executed first. To avoid that we link one chain to another. While the base chain (ddl_log_state_create) is active the secondary chain (ddl_log_state_rm) is not executed. That is: only one chain can be executed in two linked chains. The interface ddl_log_link_chains() was defined in "MDEV-22166 ddl_log_write_execute_entry() extension". Atomic info parameters in HA_CREATE_INFO Many functions in CREATE TABLE pass the same parameters. These parameters are part of table creation info and should be in HA_CREATE_INFO (or whatever). Passing parameters via single structure is much easier for adding new data and refactoring. Aria changes: - Fixed issue in Aria engine with CREATE + locked tables that data was not properly commited in some cases in case of crashes. InnoDB related changes (by Marko Mäkelä): - table_name_t::is_create_or_replace(): A new predicate to check for CREATE OR REPLACE TABLE will rename an old table to and eventually drop after creating the replacement. - dict_table_t::parse_name(): Do acquire MDL on #sql-create- names for partitioned tables. - dict_table_rename_in_cache(): On CREATE OR REPLACE TABLE ... SELECT, forget the original dict_table_t::mdl_name so that purge will acquire MDL on the #sql-create- name instead. In this way, the MDL_EXCLUSIVE that the CREATE OR REPLACE TABLE holds on the user-visible name will not unnecessarily block any purge of old history until the very end when the #sql-create- table will be dropped. - ha_innobase::delete_table(): Do not check FOREIGN KEY consistency when dropping an #sql-create- table. - row_rename_table_for_mysql(): Update SYS_FOREIGN.ID also when renaming to #sql-create- in order to avoid any duplicate key error when CREATE OR REPLACE TABLE is creating some FOREIGN KEY constraints by names that existed in the old table. Other changes: - Removed some auto variables in log.cc for better code readability. - Fixed old bug that CREATE ... SELECT would not be able to auto repair a table that is part of the SELECT. - Marked MyISAM that it does not support ROLLBACK (not required but done for better consistency with other engines). - maria_create_trn_for_mysql() does not register a new transaction handler for commits. This was needed to ensure create or replace will not end with an active transaction. - We do not get anymore warnings about "Engine not supporting atomic create" when doing a legal CREATE OR REPLACE on a table with foreign key constraints. - Updated VIDEX engine flags to disable CREATE SEQUENCE. - Removed mysql_mutex_unlock(&LOCK_gdl) / mysql_mutex_lock(&LOCK_gdl) around calls to binlog as these are unsafe. The binlog code uses global variables that needs protection from other caller. - EITS data is preserved if create or replace fails if drop_before_create_or_replace=OFF. If ON, then create or replace will drop EITS before the drop of the original table (as before). - Using CREATE OR REPLACE on a encrypted table that the user cannot decrypt will fail instead of replacing the encrypted table. The encrypted table will unchanged. Known issues: - One cannot use create or replace on an InnoDB tables that has foreign key references point to it - CREATE OR REPLACE TEMPORARY table is not full atomic. Any conflicting table will always be dropped before creating a new one. (Old behaviour). Bug fixes related to this MDEV: MDEV-36435 Assertion failure in finalize_locked_tables() MDEV-36439 Assertion `thd_arg->lex->sql_command != SQLCOM_CREATE_SEQUENCE... MDEV-36498 Failed CoR in non-atomic mode no longer generates DROP in RBR... MDEV-36508 Temporary files #sql-create-....frm occasionally stay after crash recovery MDEV-38479 Crash in CREATE OR REPLACE SEQUENCE when new sequence cannot be created MDEV-36497 Assertion failure after atomic CoR with Aria under lock in transactional context MDEV-36501 EITS data is lost after failed attempt to CREATE OR REPLACE table MDEV-36493 Atomic CREATE OR REPLACE ... SELECT blocks InnoDB purge MDEV-39367 MSAN/valgrind errors in temp_file_size_cb_func, main.tmp_space_usage fails MDEV-39446 Atomic CREATE OR REPLACE fails if a table cannot be decrypted MDEV-40776 Atomic CREATE OR REPLACE silently breaks the foreign key MDEV-40765 Assertion `"unexpected references" == 0' failed upon failing CREATE OR REPLACE MDEV-40994 Atomic CoR: Transaction not registered for MariaDB 2PC, but transaction is active Reverted commits: - MDEV-36685 "CREATE-SELECT may lose in binlog side-effects of stored-routine" as it did not take into account that it safe to clear binlogs if the created table is non transactional and there are no other non transactional tables used. This was done because it caused extra logging when it is not needed (not using any non transactional tables) and it also did not solve side effects when using statement based loggging. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-25529 Timestamp_string for printing timestamps Add a Timestamp_string helper that formats a Timestamp in the session time zone for use in warning and error messages, plus THD::timestamp_to_string() and a Temporal_hybrid constructor behind it. Reuse the existing Timestamp type rather than introducing a new one. Use it when printing STARTS and the history range. Rename the error symbol ER_PART_STARTS_BEYOND_INTERVAL to WARN_VERS_STARTS_BEYOND_INTERVAL and extend its message: it now takes the STARTS timestamp and the query timestamp as arguments, so the number of format arguments changes from one to three. The error number (4164) is unchanged, so this is not an ABI break, but libmariadb still exports the old symbol name for that number until its submodule is updated. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Monty
monty@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Added --debug-dbug option to mysqltest.cc This was to get rid of warnings when using mtr --debug |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-25529 set_up_default_partitions() ER_OUT_OF_RESOURCES error | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alexander Barkov
bar@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41246 "Illegal mix of collations" on the mysql.user view This SQL script failed: SET NAMES latin1 COLLATE latin1_swedish_ci; CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1; SET NAMES big5 COLLATE big5_chinese_ci; SELECT * FROM v1 WHERE c1='y'; with the following error: ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '=' Note, latin1_swedish_ci and big5_chinese_ci are used here as examples. The error also happened with different collation combinations. Fix main idea: If two collations have equal comparison rules (known as "tailoring") on a given character repertoire, like latin1_swedish_ci and big5_chinese_ci on ASCII letters, then the "Illegal mix of collations" error can be avoided in a comparison operator. We can choose any of the sides as the operation effective collation - the result will be equal. This optimization is not applied when at least one side has an explicit COLLATE clause. Two explicit COLLATE clauses in one comparison are already illegal when the character sets are the same, so for consistency this stays illegal when the character sets differ too. Most important details: - Splitting enum_repertoire_t into smaller subsets, for better repertoire granularity. A variable holding a repertoire value can now have multiple MY_REPERTOIRE_XXX flags set. This patch implements detecting tailoring equality on these repertoires: * MY_REPERTOIRE_ASCII_DIGITS - [0..9] * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'. It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and MY_REPERTOIRE_ASCII_DOT. They were a part of MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed (223 and 255). * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9]. * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore * MY_REPERTOIRE_ASCII - the entire range U+0000..U+007F - Adding a new virtual function "tailoring" in my_collation_handler_st. It returns the tailoring on the given repertoire for the given collation. If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0}, it means the illegal mix optimization cannot be used for this collation on the given repertoire. If these calls: tr1= cs1->coll->tailoring(cs1, some_repertoire); tr2= cs2->coll->tailoring(cs2, some_repertoire); return both non-NULL results and tr1.str==tr2.str, then these collations are equal on the given repertoire and are mutually replaceable for a comparison operator, so "Illegal mix of collations" can be avoided. - Tailoring strings are shared constants my_tailoring_str_* (strings/strings_def.h). Tailorings are compared by the string pointer, so tailoring() implementations return strings only from these constants. The strings do not depend on PAD/NOPAD. - my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit character sets no longer stops at the first bad byte sequence with the repertoire found so far, which could be MY_REPERTOIRE_NONE (compatible with everything). It returns MY_REPERTOIRE_EXTENDED for such strings. A unit test was added into unittest/strings. - Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set can store all ASCII characters (it is false for swe7). It is used in Item_func_conv_charset to calculate "safe", and in left_is_algorithmically_simpler() to prefer other character sets to those where ASCII strings cannot be converted safely. - DBUG_ASSERTs were added into mysys/charset.c, to check that the "tailoring" method is set in collation handlers when a collation is added: compiled-in, from ctype-extra.c, or loaded from Index.xml. - Adding a new function my_collations_equal_on_repertoire(cs1, cs2, repertoire) in strings/. It tells if two collations are equal on the repertoire, taking tailoring() results and PAD/NOPAD and the UCA version into account. Callers do not compare tailoring() results directly, so the way tailorings are compared (currently by the pointer to the shared constant string) is private to strings/ and can be extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30. - Adding a new method DTCollation::aggregate_by_tailoring(). Collations with different MY_CS_NOPAD are never compatible. UCA collations of different versions are not compatible on the repertoire with ASCII punctuation, as the UCA version affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT and CIRCUMFLEX ACCENT after PERCENT SIGN). - DTCollation::aggregate() now makes the repertoire of the result cover the repertoires of both sides in all successful branches. Before, set(dt) replaced the repertoire with the repertoire of the winner side only, so the result could have a too narrow repertoire, e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the value '1.5' has a punctuation character. With the more granular repertoires this could wrongly allow mixing collations which are equal on ALNUM but different on punctuation. - Adding a new flag MY_COLL_ALLOW_BY_TAILORING. It indicates to DTCollation::aggregate() that the illegal mix optimization by repertoire can be used in the given context. MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING. Note, only comparison operators pass this flag. Functions returning a string result do not pass this flag, because in operations like CONCAT(a,b) we still need to evaluate precisely the collation of the result - we cannot just choose a collation of one of the sides (even if they are compatible on the given repertoire). This also applies to ExtractValue() and UpdateXML(), which now aggregate their arguments without this flag. - As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits rather than a single bit, the way to detect "is only ASCII" repertoires has changed in the code. For example: // repertoire *IS* ascii if (repertoire == MY_REPERTOIRE_ASCII) has changed in multiple places in the code to // repertoire *HAS* only ascii characters if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII)) The new function my_repertoire_is_subset_of() in m_ctype.h is used for this purpose in C code. In C++ code DTCollation::repertoire_is_subset_of() is used, e.g.: if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII)) - Repertoire of some expressions was adjusted to the new meaning: * Item_null now has MY_REPERTOIRE_NONE (was ASCII). * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30 (was EXTENDED). * Lex_string_with_metadata_st::repertoire(cs) now scans the string contents to detect the actual repertoire. * HEX() now has MY_REPERTOIRE_ASCII_ALNUM. * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale which is actually used: the explicit third argument, or @@lc_time_names. It was always @@lc_time_names before. A non-constant locale argument is assumed to be non-ASCII. * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY, JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra characters they put into the result (quotes, separators, padding, brackets). * LOWER() and UPPER() add the letters of both cases to the repertoire, because the result can contain letters of the opposite case. In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an ASCII letter can be converted to a non-ASCII one (I -> dotless i). Unicode collations are detected by the new member casefold_info_st::can_convert_to_non_ascii_on_casefolding, simple 8bit collations by to_lower['I'] and to_upper['i']. - Lex_string_with_metadata_st::repertoire(cs) now scans a string literal for any byte oriented character set (mbminlen==1), not only for ASCII based ones. It was assumed to have the full repertoire for swe7, which is not ASCII based. The scan is done by Unicode code points, so swe7 strings with digits and letters get a narrow repertoire, while e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII. Therefore swe7_swedish_ci can be mixed with other collations now, in the same way as other collations with MY_CS_IDENT_CASEUP_CI. The test func_time was adjusted: a swe7 literal ' ull' with ASCII characters is now safe to convert, the error is still tested for '[ull'. - cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS, as it sorts the minus and the dot differently. Both were verified with ctype_digits_verify. - The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL), so my_charset_is_ascii_based() is true for tis620. It was false for tis620, although tis620 is ASCII compatible. Therefore tis620 string literals were always considered to have a non-ASCII repertoire, and the repertoire optimizations did not work for them. Now ASCII-only tis620 literals are detected as ASCII. This changes the result of subselect_extra_no_semijoin from an error to success. - tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case and then compare by the code point on the entire ASCII range (the Thai specific rules affect only non-ASCII characters). A new function my_tailoring_ascii_casedn_ci() returns tailorings for such collations on ALNUM, IDENT and ASCII. For example, the underscore sorts before the letters, unlike in the caseup tailorings. - New flags were added for CHARSET_INFO::state * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the ASCII repertoire, and converts lower case letters to upper case ones, so the underscore sorts after the letters (with an upper to lower case conversion it would sort before the letters). * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the IDENT repertoire (but can have irregularities say on punctuation). * MY_CS_ASCII_STD_UCA - for UCA collations. It means that a UCA collation does not reorder ASCII characters. * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits 0..9 are sorted in the code point order and are not tailored, and, for PAD collations, the space sorts before the digits. Strings of digits are compared in the same way in all such collations, even if they have irregularities on letters (latin7, cp866, czech, danish, turkish UCA, etc.), or differ by case sensitivity. It is detected from sort_order for 8bit collations, and from the tailoring rules for UCA collations. Collations with MY_CS_DIGITS_STD still must have the same PAD/NOPAD attribute. * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means MY_CS_DIGITS_STD, and additionally the minus sorts before the dot, which sorts before the digits (and the space sorts before the minus in PAD collations). Strings like '-10.5' are compared in the same way in all such collations. It is detected from sort_order for simple 8bit collations (my_coll_init_simple()), and from the tailoring rules for UCA collations. For the multi-byte collations big5_chinese_ci, gbk_chinese_ci, gb2312_chinese_ci and sjis_japanese_ci it is declared by my_tailoring_ident_caseup_ci_generic(). The flags MY_CS_DIGITS_STD and MY_CS_ASCII_MINUS_DOT_DIGITS are stored in ctype-extra.c for compiled collations (conf_to_src.c prints them), and the detection is skipped for a collation which already has any of the repertoire flags (MY_CS_REPERTOIRE_FLAGS). - strings/conf_to_src.c was modified to detect and print MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags. - strings/ctype-extra.c was regenerated with new flags. - 8bit collations are not given MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort before the digits. A shorter string is padded with spaces for comparison, so the position of the space matters even for strings without spaces ("ab" is compared to "abc" as "ab " to "abc"). A test collation latin1_space_after_letters_ci was added to mysql-test/std_data/ldml/ for this. ctype-extra.c does not change. - The initializer of my_collation_cs_handler in ctype-utf8.c (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match my_collation_handler_st: the missing get_id and get_collation_name members were added, so that eq_collation was not initialized in their place, and the new member "tailoring" is now set to my_tailoring_none. - Results of the existing tests func_str, func_test, ps, view and subselect_sj changed from "Illegal mix of collations" to success. In these tests ASCII-only strings were compared using collations with equal tailoring on ASCII, e.g. latin1_swedish_ci with latin2_general_ci (func_str, ps), koi8r_general_ci with latin1_swedish_ci (func_test), latin1_general_ci with latin1_swedish_ci on letters and digits (view), cp932_japanese_ci with latin1_swedish_ci (subselect_sj). Such comparisons are now allowed. Cases with collations that are still incompatible on ASCII, e.g. latin7_general_ci, were added to the tests and still fail. - Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/. They display a tailoring by collation name and repertoire as returned by: cs->coll->tailoring(cs, some_repertoire) A dynamically linked plugin function collation_tailoring() was added for the purpose of these tests. Tailorings do not depend on PAD/NOPAD and on the UCA version, so the function shows "[nopad]" for NOPAD collations, and "[version X.Y.Z]" for UCA collations on the ASCII repertoire. The version is formatted by a new class UCAVersion, which uses a new method CharBuffer::append_uint8(). - Adding the test plugin/func_test/mysql-test/func_test/ ctype_caseup_ci_verify.test. It checks the order of ASCII characters for all collations declared "caseup CI" on the ASCII or IDENT range. - Adding the test plugin/func_test/mysql-test/func_test/ collation_tailoring_args.test. It checks the arguments of the plugin function collation_tailoring(): the repertoire passed by name or by number, and the NULL results for unknown names. - Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test. It displays the tailorings of collations loaded at runtime from Index.xml: simple 8bit collations (flags detected from the weights) and Unicode collations (LDML rules which do or do not reorder ASCII). - Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test They display two-dimensional charts showing which collations are compatible on which repertoires. Tests for DATE_FORMAT() with an explicit locale were added to ctype_big5_tailoring.test, ctype_cp932_tailoring.test and ctype_latin5_tailoring.test, which also has the chart for latin5. - Adding a number of MTR tests in the form of the originally reported script for various collations: SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...'; The repeated blocks of queries are shared through mysql-test/include/ ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc. - Adding tests to ctype_cp932.test checking that ExtractValue() and UpdateXML() still raise "Illegal mix of collations". |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
iff-sal
iffathdev@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40296: SELECT @@GLOBAL.replicate_do_db ignores default_master_connection When querying replication filter system variables without specifying a connection name qualifier (e.g., SELECT @@GLOBAL.replicate_do_db), Sys_var_rpl_filter::global_value_ptr() previously passed the unspecified base_name directly to get_master_info(), which resolved to the default connection (empty_clex_str) rather than respecting @@SESSION.default_master_connection. Fix by falling back to thd->variables.default_master_connection when base_name->str is NULL in Sys_var_rpl_filter::global_value_ptr(). Add regression test in suite/multi_source/rpl_filter_default_master.test covering explicit named connection, unspecified connection falling back to default, and explicit unnamed connection with empty backticks. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
yuchen.pei@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Allow a covering index scan for a locking SELECT under a full-scan table lock Follow-up to MDEV-24813 (innodb_table_lock_on_full_scan) and MDEV-40805. ha_innobase::build_template() retrieves the whole clustered index record whenever select_lock_type is LOCK_X, and row_search_with_covering_prefix() refuses the covering-index optimisation for the same reason. As a result a covering secondary index scan such as SELECT sec, id FROM t1 FOR UPDATE visits the clustered index once per scanned row, while the same scan with LOCK IN SHARE MODE stays inside the secondary index. On a one million row table that is a million extra clustered index lookups. Only part of that work is inherent to LOCK_X. UPDATE and DELETE do need the clustered index record, because they are going to write it. A plain locking SELECT does not: it visits the clustered index only to place the per-row exclusive lock that gives SELECT ... FOR UPDATE its row level exclusivity. When innodb_table_lock_on_full_scan made us take a table level LOCK_X for the whole scan, that per-row lock is redundant, because the table lock is already mutually exclusive with any other transaction's LOCK_IX, and hence with any record lock or implicit exclusive lock in the table. Introduce row_prebuilt_t::full_scan_covering_read, set in ha_innobase::extra_opt() next to full_table_scan and only when the statement is a plain SELECT outside the HANDLER interface, and skip the LOCK_X restriction in both places when it is set. The flag is never set unless full_table_scan is set, so row level locking behaviour is unchanged when innodb_table_lock_on_full_scan is off. Whether a column outside the scanned index is needed is still decided by the existing logic in build_template(), so SELECT pad ... FOR UPDATE continues to read the clustered index. Co-Authored-By: Claude Opus 5 <[email protected]> fixup: keep the clustered index visit under snapshot isolation innodb.lock_isolation 'table_lock' failed in the MDEV-33802 section: SELECT * FROM t FORCE INDEX (b) FOR UPDATE succeeded where ER_CHECKREAD was expected. On t(a INT PRIMARY KEY, b INT UNIQUE) the secondary index b stores (b, a), so SELECT * is covered by it, and the previous commit let a locking SELECT take the covering path under a full-scan table LOCK_X. Placing the record lock is not the only thing the clustered index visit does. With innodb_snapshot_isolation and a read view already open, lock_clust_rec_read_check_and_lock() also reads the clustered record's DB_TRX_ID and returns DB_RECORD_CHANGED when the read view cannot see it. Skipping the clustered index skips that check, so the statement silently locked a row it should have refused. The table-level lock does not substitute for the check. It gives exclusivity against concurrent transactions, whereas this reports a change that committed before the lock was taken and is invisible to an older read view. DB_TRX_ID is only stored in the clustered index record, so the check cannot be answered from a secondary index record alone. Keep visiting the clustered index whenever snapshot isolation is active and a read view is open. A plain locking SELECT that opens no read view, which is the case the optimisation targets, is unaffected. Co-Authored-By: Claude Opus 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-25529 cleanup for vers_set_starts() and starts_clause | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Marko Mäkelä
marko.makela@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
fixup! 276b6dd3a08fec92e474963e8defecd22bc092ad error C2886: 'tpool::pwrite': symbol cannot be used in a member using-declaration |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-25529 Auto-create: Pre-existing historical data is not partitioned as specified by ALTER Adds logic into prep_alter_part_table() for AUTO to check the history range (vers_get_history_range()) and based on (max_ts - min_ts) difference compute the number of created partitions and set STARTS value to round down min_ts value (vers_set_starts()) if it was not specified by user or if the user specified it incorrectly. In the latter case it will print warning about wrongly specified user value. In case of fast ALTER TABLE, f.ex. when partitioning already exists, the above logic is ignored unless FORCE clause is specified. When user specifies partition list explicitly the above logic is ignored even with FORCE clause. vers_get_history_range() detects if the index can be used for row_end min/max stats and if so it gets it with ha_index_first() and HA_READ_BEFORE_KEY (as it must ignore current data). Otherwise it does table scan to read the stats. There is test_mdev-25529 debug keyword to check the both and compare results. A warning is printed if the algorithm uses slow scan. Static key_cmp was renamed to key_eq to resolve compilation after key.h was included as key_cmp was already declared there. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-25529 converted COMBINE macro to interval2usec inline function | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Georgi (Joro) Kodinov
joro@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| more doxygen formatting added. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Kristian Nielsen
knielsen@knielsen-hq.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41217: Inefficient memory management by parallel slave Draft patch / proof-of-concept. This patch splits the rpl_group_info into two parts. A small part that contains only the information necessary in the queue of events sent from the SQL driver thread to parallel replication worker threads is made separate and pointed to by rpl_group_info::q. This greatly reduces the memory requirements for the queued event groups, especially when the slave is lagging and there are a lot of event groups queued (eg. due to large --slave-parallel-max-queued). Now the main rpl_group_info object only exists once per worker thread (and once in the SQL driver thread). Signed-off-by: Kristian Nielsen <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alexander Barkov
bar@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41246 "Illegal mix of collations" on the mysql.user view This SQL script failed: SET NAMES latin1 COLLATE latin1_swedish_ci; CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1; SET NAMES big5 COLLATE big5_chinese_ci; SELECT * FROM v1 WHERE c1='y'; with the following error: ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '=' Note, latin1_swedish_ci and big5_chinese_ci are used here as examples. The error also happened with different collation combinations. Fix main idea: If two collations have equal comparison rules (known as "tailoring") on a given character repertoire, like latin1_swedish_ci and big5_chinese_ci on ASCII letters, then the "Illegal mix of collations" error can be avoided in a comparison operator. We can choose any of the sides as the operation effective collation - the result will be equal. This optimization is not applied when at least one side has an explicit COLLATE clause. Two explicit COLLATE clauses in one comparison are already illegal when the character sets are the same, so for consistency this stays illegal when the character sets differ too. Most important details: - Splitting enum_repertoire_t into smaller subsets, for better repertoire granularity. A variable holding a repertoire value can now have multiple MY_REPERTOIRE_XXX flags set. This patch implements detecting tailoring equality on these repertoires: * MY_REPERTOIRE_ASCII_DIGITS - [0..9] * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'. It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and MY_REPERTOIRE_ASCII_DOT. They were a part of MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed (223 and 255). * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9]. * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore * MY_REPERTOIRE_ASCII - the entire range U+0000..U+007F - Adding a new virtual function "tailoring" in my_collation_handler_st. It returns the tailoring on the given repertoire for the given collation. If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0}, it means the illegal mix optimization cannot be used for this collation on the given repertoire. If these calls: tr1= cs1->coll->tailoring(cs1, some_repertoire); tr2= cs2->coll->tailoring(cs2, some_repertoire); return both non-NULL results and tr1.str==tr2.str, then these collations are equal on the given repertoire and are mutually replaceable for a comparison operator, so "Illegal mix of collations" can be avoided. - Tailoring strings are shared constants my_tailoring_str_* (strings/strings_def.h). Tailorings are compared by the string pointer, so tailoring() implementations return strings only from these constants. The strings do not depend on PAD/NOPAD. - my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit character sets no longer stops at the first bad byte sequence with the repertoire found so far, which could be MY_REPERTOIRE_NONE (compatible with everything). It returns MY_REPERTOIRE_EXTENDED for such strings. A unit test was added into unittest/strings. - Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set can store all ASCII characters (it is false for swe7). It is used in Item_func_conv_charset to calculate "safe", and in left_is_algorithmically_simpler() to prefer other character sets to those where ASCII strings cannot be converted safely. - DBUG_ASSERTs were added into mysys/charset.c, to check that the "tailoring" method is set in collation handlers when a collation is added: compiled-in, from ctype-extra.c, or loaded from Index.xml. - Adding a new function my_collations_equal_on_repertoire(cs1, cs2, repertoire) in strings/. It tells if two collations are equal on the repertoire, taking tailoring() results and PAD/NOPAD and the UCA version into account. Callers do not compare tailoring() results directly, so the way tailorings are compared (currently by the pointer to the shared constant string) is private to strings/ and can be extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30. - Adding a new method DTCollation::aggregate_by_tailoring(). Collations with different MY_CS_NOPAD are never compatible. UCA collations of different versions are not compatible on the repertoire with ASCII punctuation, as the UCA version affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT and CIRCUMFLEX ACCENT after PERCENT SIGN). - DTCollation::aggregate() now makes the repertoire of the result cover the repertoires of both sides in all successful branches. Before, set(dt) replaced the repertoire with the repertoire of the winner side only, so the result could have a too narrow repertoire, e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the value '1.5' has a punctuation character. With the more granular repertoires this could wrongly allow mixing collations which are equal on ALNUM but different on punctuation. - Adding a new flag MY_COLL_ALLOW_BY_TAILORING. It indicates to DTCollation::aggregate() that the illegal mix optimization by repertoire can be used in the given context. MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING. Note, only comparison operators pass this flag. Functions returning a string result do not pass this flag, because in operations like CONCAT(a,b) we still need to evaluate precisely the collation of the result - we cannot just choose a collation of one of the sides (even if they are compatible on the given repertoire). This also applies to ExtractValue() and UpdateXML(), which now aggregate their arguments without this flag. - As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits rather than a single bit, the way to detect "is only ASCII" repertoires has changed in the code. For example: // repertoire *IS* ascii if (repertoire == MY_REPERTOIRE_ASCII) has changed in multiple places in the code to // repertoire *HAS* only ascii characters if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII)) The new function my_repertoire_is_subset_of() in m_ctype.h is used for this purpose in C code. In C++ code DTCollation::repertoire_is_subset_of() is used, e.g.: if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII)) - Repertoire of some expressions was adjusted to the new meaning: * Item_null now has MY_REPERTOIRE_NONE (was ASCII). * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30 (was EXTENDED). * Lex_string_with_metadata_st::repertoire(cs) now scans the string contents to detect the actual repertoire. * HEX() now has MY_REPERTOIRE_ASCII_ALNUM. * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale which is actually used: the explicit third argument, or @@lc_time_names. It was always @@lc_time_names before. A non-constant locale argument is assumed to be non-ASCII. * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY, JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra characters they put into the result (quotes, separators, padding, brackets). * LOWER() and UPPER() add the letters of both cases to the repertoire, because the result can contain letters of the opposite case. In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an ASCII letter can be converted to a non-ASCII one (I -> dotless i). Unicode collations are detected by the new member casefold_info_st::can_convert_to_non_ascii_on_casefolding, simple 8bit collations by to_lower['I'] and to_upper['i']. - Lex_string_with_metadata_st::repertoire(cs) now scans a string literal for any byte oriented character set (mbminlen==1), not only for ASCII based ones. It was assumed to have the full repertoire for swe7, which is not ASCII based. The scan is done by Unicode code points, so swe7 strings with digits and letters get a narrow repertoire, while e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII. Therefore swe7_swedish_ci can be mixed with other collations now, in the same way as other collations with MY_CS_IDENT_CASEUP_CI. The test func_time was adjusted: a swe7 literal ' ull' with ASCII characters is now safe to convert, the error is still tested for '[ull'. - cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS, as it sorts the minus and the dot differently. Both were verified with ctype_digits_verify. - The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL), so my_charset_is_ascii_based() is true for tis620. It was false for tis620, although tis620 is ASCII compatible. Therefore tis620 string literals were always considered to have a non-ASCII repertoire, and the repertoire optimizations did not work for them. Now ASCII-only tis620 literals are detected as ASCII. This changes the result of subselect_extra_no_semijoin from an error to success. - tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case and then compare by the code point on the entire ASCII range (the Thai specific rules affect only non-ASCII characters). A new function my_tailoring_ascii_casedn_ci() returns tailorings for such collations on ALNUM, IDENT and ASCII. For example, the underscore sorts before the letters, unlike in the caseup tailorings. - New flags were added for CHARSET_INFO::state * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the ASCII repertoire, and converts lower case letters to upper case ones, so the underscore sorts after the letters (with an upper to lower case conversion it would sort before the letters). * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the IDENT repertoire (but can have irregularities say on punctuation). * MY_CS_ASCII_STD_UCA - for UCA collations. It means that a UCA collation does not reorder ASCII characters. * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits 0..9 are sorted in the code point order and are not tailored, and, for PAD collations, the space sorts before the digits. Strings of digits are compared in the same way in all such collations, even if they have irregularities on letters (latin7, cp866, czech, danish, turkish UCA, etc.), or differ by case sensitivity. It is detected from sort_order for 8bit collations, and from the tailoring rules for UCA collations. Collations with MY_CS_DIGITS_STD still must have the same PAD/NOPAD attribute. * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means MY_CS_DIGITS_STD, and additionally the minus sorts before the dot, which sorts before the digits (and the space sorts before the minus in PAD collations). Strings like '-10.5' are compared in the same way in all such collations. It is detected from sort_order for simple 8bit collations (my_coll_init_simple()), and from the tailoring rules for UCA collations. For the multi-byte collations big5_chinese_ci, gbk_chinese_ci, gb2312_chinese_ci and sjis_japanese_ci it is declared by my_tailoring_ident_caseup_ci_generic(). The flags MY_CS_DIGITS_STD and MY_CS_ASCII_MINUS_DOT_DIGITS are stored in ctype-extra.c for compiled collations (conf_to_src.c prints them), and the detection is skipped for a collation which already has any of the repertoire flags (MY_CS_REPERTOIRE_FLAGS). - strings/conf_to_src.c was modified to detect and print MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags. - strings/ctype-extra.c was regenerated with new flags. - 8bit collations are not given MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort before the digits. A shorter string is padded with spaces for comparison, so the position of the space matters even for strings without spaces ("ab" is compared to "abc" as "ab " to "abc"). A test collation latin1_space_after_letters_ci was added to mysql-test/std_data/ldml/ for this. ctype-extra.c does not change. - The initializer of my_collation_cs_handler in ctype-utf8.c (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match my_collation_handler_st: the missing get_id and get_collation_name members were added, so that eq_collation was not initialized in their place, and the new member "tailoring" is now set to my_tailoring_none. - Results of the existing tests func_str, func_test, ps, view and subselect_sj changed from "Illegal mix of collations" to success. In these tests ASCII-only strings were compared using collations with equal tailoring on ASCII, e.g. latin1_swedish_ci with latin2_general_ci (func_str, ps), koi8r_general_ci with latin1_swedish_ci (func_test), latin1_general_ci with latin1_swedish_ci on letters and digits (view), cp932_japanese_ci with latin1_swedish_ci (subselect_sj). Such comparisons are now allowed. Cases with collations that are still incompatible on ASCII, e.g. latin7_general_ci, were added to the tests and still fail. - Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/. They display a tailoring by collation name and repertoire as returned by: cs->coll->tailoring(cs, some_repertoire) A dynamically linked plugin function collation_tailoring() was added for the purpose of these tests. Tailorings do not depend on PAD/NOPAD and on the UCA version, so the function shows "[nopad]" for NOPAD collations, and "[version X.Y.Z]" for UCA collations on the ASCII repertoire. The version is formatted by a new class UCAVersion, which uses a new method CharBuffer::append_uint8(). - Adding the test plugin/func_test/mysql-test/func_test/ ctype_caseup_ci_verify.test. It checks the order of ASCII characters for all collations declared "caseup CI" on the ASCII or IDENT range. - Adding the test plugin/func_test/mysql-test/func_test/ collation_tailoring_args.test. It checks the arguments of the plugin function collation_tailoring(): the repertoire passed by name or by number, and the NULL results for unknown names. - Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test. It displays the tailorings of collations loaded at runtime from Index.xml: simple 8bit collations (flags detected from the weights) and Unicode collations (LDML rules which do or do not reorder ASCII). - Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test They display two-dimensional charts showing which collations are compatible on which repertoires. Tests for DATE_FORMAT() with an explicit locale were added to ctype_big5_tailoring.test, ctype_cp932_tailoring.test and ctype_latin5_tailoring.test, which also has the chart for latin5. - Adding a number of MTR tests in the form of the originally reported script for various collations: SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...'; The repeated blocks of queries are shared through mysql-test/include/ ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc. - Adding tests to ctype_cp932.test checking that ExtractValue() and UpdateXML() still raise "Illegal mix of collations". |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Khaled Riyad
khaled57.dev@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41140 mariadb-server.postinst chown dereferences symlinks The postinst and preinst scripts followed symlinks when fixing ownership under /var/log/mysql and /var/lib/mysql. Use -h and stop find from descending so symlinks there are no longer dereferenced. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Oleksandr Byelkin
sanja@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-32401 expression cache lead to crash if table of wrong type created and the cache switched off 1) take into account TMP_TABLE_ALL_COLUMNS when we are modifying agg_item->result_field 2) remove unused now "bool materialized_subquery;" Co-Authored-By: Claude Opus 5.5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Pekka Lampio
pekka.lampio@galeracluster.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40018: SHOW SLAVE STATUS allocates ~1.7 MB per call due to oversized VARCHAR columns SHOW SLAVE STATUS / SHOW ALL SLAVES STATUS / INFORMATION_SCHEMA.SLAVE_STATUS are all served from the ST_FIELD_INFO array slave_status_info[]. Thirteen of its columns (Replicate_Do_DB, Replicate_Ignore_DB, Replicate_Do_Table, Replicate_Ignore_Table, Replicate_Wild_Do_Table, Replicate_Wild_Ignore_Table, Replicate_Ignore_Server_Ids, Gtid_IO_Pos, Replicate_Do_Domain_Ids, Replicate_Ignore_Domain_Ids, Slave_SQL_Running_State, Replicate_Rewrite_DB, and Gtid_Slave_Pos) used the no-argument Varchar() helper, which defaults to MAX_FIELD_VARCHARLENGTH/3 characters. With utf8mb3, that reserves 65532 bytes per column in the temporary table's record buffer regardless of the actual value size, and the record is allocated twice, accounting for the reported ~1.7 MB allocated on every call. Change these columns to Blob(MAX_FIELD_VARCHARLENGTH), which keeps the same maximum capacity but stores only a pointer and length in the record. Values, NULL-ability, and column order are unchanged; the client-visible column type for these columns changes from VARCHAR to BLOB. mysql-test/suite/funcs_1/r/is_columns_is.result is updated to match, since it asserts on INFORMATION_SCHEMA.COLUMNS metadata for these columns. Assisted-By: Claude AI |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41351 Allow a covering index scan for a SELECT ... FOR UPDATE under a full-scan table lock (TODO: improve commit message) MDEV-24813 introduced a switch innodb_table_lock_on_full_scan that places a table lock on an innodb table for some unconditional SELECT statements. With a full table lock, and when doing a covering scan, there is no need to acquire clustered locks. This patch does so for X-locks. TODO: add code comments and tests Co-Authored-By: Claude Opus 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Sergei Petrunia
sergey@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-40553: include/opt_context_save_in_var.inc can't handle quotes I_S.OPTIMIZER_CONTEXT has this statement: set @opt_context='CONTEXT'; -- denotes CONTEXT-STMT And opt_context_save_in_var.inc just took the CONTEXT and put it in a variable: SET @opt_context= (select REGEXP_SUBSTR(...) from i_s.optimizer_context); This ignored the fact that the "CONTEXT" is escaped for the SQL parser. Fix this by extracting the entire CONTEXT-STMT and running it. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41351 Allow a covering index scan for a SELECT ... FOR UPDATE under a full-scan table lock (TODO: improve commit message) MDEV-24813 introduced a switch innodb_table_lock_on_full_scan that places a table lock on an innodb table for some unconditional SELECT statements. With a full table lock, and when doing a covering scan, there is no need to acquire clustered locks. This patch does so for X-locks. TODO: add code comments Co-Authored-By: Claude Opus 5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Kristian Nielsen
knielsen@knielsen-hq.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Small docs update for binlog-in-engine Signed-off-by: Kristian Nielsen <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Marko Mäkelä
marko.makela@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
fixup! 1b60442fec1e7c5dfa25915d0fa7fd8aa6c4ea88 Guarantee to copy the log file in log_sys.write_size blocks |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Monty
monty@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
ha_table_exists() cleanup and improvement This is part of MDEV-25292 Atomic CREATE OR REPLACE TABLE. Removed default values for arguments, added flags argument to specify filename flags (FN_TO_IS_TMP, FN_FROM_IS_TMP) and forward the flag to build_table_name(). Original patch from: Aleksey Midenkov <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Monty
monty@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MDEV-31138 Enable the test spider/bugfix.mdev_29676 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Oleksandr Byelkin
sanja@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41099 HANDLER READ missing column privilege check HANDLER OPEN accepted column-level SELECT grants as table access, but HANDLER READ returns all columns and never checked column privileges. Reopen of a handler table (after FLUSH/ALTER) did no checks at all. Require SELECT on every column (as for SELECT *) when the user has no table-level SELECT, and redo the privilege checks on reopen. Co-Authored-By: Claude Opus 5.5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Monty
monty@mariadb.org |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-23298 Assertion `table_list->prelocking_placeholder == TABLE_LIST::PRELOCK_NONE' failed in check_lock_and_start_stmt on CREATE OR REPLACE TABLE Fixed by removing wrong assert Review: Sanja Byelkin |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Alexander Barkov
bar@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41246 "Illegal mix of collations" on the mysql.user view This SQL script failed: SET NAMES latin1 COLLATE latin1_swedish_ci; CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1; SET NAMES big5 COLLATE big5_chinese_ci; SELECT * FROM v1 WHERE c1='y'; with the following error: ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '=' Note, latin1_swedish_ci and big5_chinese_ci are used here as examples. The error also happened with different collation combinations. Fix main idea: If two collations have equal comparison rules (known as "tailoring") on a given character repertoire, like latin1_swedish_ci and big5_chinese_ci on ASCII letters, then the "Illegal mix of collations" error can be avoided in a comparison operator. We can choose any of the sides as the operation effective collation - the result will be equal. This optimization is not applied when at least one side has an explicit COLLATE clause. Two explicit COLLATE clauses in one comparison are already illegal when the character sets are the same, so for consistency this stays illegal when the character sets differ too. Most important details: - Splitting enum_repertoire_t into smaller subsets, for better repertoire granularity. A variable holding a repertoire value can now have multiple MY_REPERTOIRE_XXX flags set. This patch implements detecting tailoring equality on these repertoires: * MY_REPERTOIRE_ASCII_DIGITS - [0..9] * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'. It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and MY_REPERTOIRE_ASCII_DOT. They were a part of MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed (223 and 255). * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9]. * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore * MY_REPERTOIRE_ASCII - the entire range U+0000..U+007F - Adding a new virtual function "tailoring" in my_collation_handler_st. It returns the tailoring on the given repertoire for the given collation. If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0}, it means the illegal mix optimization cannot be used for this collation on the given repertoire. If these calls: tr1= cs1->coll->tailoring(cs1, some_repertoire); tr2= cs2->coll->tailoring(cs2, some_repertoire); return both non-NULL results and tr1.str==tr2.str, then these collations are equal on the given repertoire and are mutually replaceable for a comparison operator, so "Illegal mix of collations" can be avoided. - Tailoring strings are shared constants my_tailoring_str_* (strings/strings_def.h). Tailorings are compared by the string pointer, so tailoring() implementations return strings only from these constants. The strings do not depend on PAD/NOPAD. - my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit character sets no longer stops at the first bad byte sequence with the repertoire found so far, which could be MY_REPERTOIRE_NONE (compatible with everything). It returns MY_REPERTOIRE_EXTENDED for such strings. A unit test was added into unittest/strings. - Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set can store all ASCII characters (it is false for swe7). It is used in Item_func_conv_charset to calculate "safe", and in left_is_algorithmically_simpler() to prefer other character sets to those where ASCII strings cannot be converted safely. - DBUG_ASSERTs were added into mysys/charset.c, to check that the "tailoring" method is set in collation handlers when a collation is added: compiled-in, from ctype-extra.c, or loaded from Index.xml. - Adding a new function my_collations_equal_on_repertoire(cs1, cs2, repertoire) in strings/. It tells if two collations are equal on the repertoire, taking tailoring() results and PAD/NOPAD and the UCA version into account. Callers do not compare tailoring() results directly, so the way tailorings are compared (currently by the pointer to the shared constant string) is private to strings/ and can be extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30. - Adding a new method DTCollation::aggregate_by_tailoring(). Collations with different MY_CS_NOPAD are never compatible. UCA collations of different versions are not compatible on the repertoire with ASCII punctuation, as the UCA version affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT and CIRCUMFLEX ACCENT after PERCENT SIGN). - DTCollation::aggregate() now makes the repertoire of the result cover the repertoires of both sides in all successful branches. Before, set(dt) replaced the repertoire with the repertoire of the winner side only, so the result could have a too narrow repertoire, e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the value '1.5' has a punctuation character. With the more granular repertoires this could wrongly allow mixing collations which are equal on ALNUM but different on punctuation. - Adding a new flag MY_COLL_ALLOW_BY_TAILORING. It indicates to DTCollation::aggregate() that the illegal mix optimization by repertoire can be used in the given context. MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING. Note, only comparison operators pass this flag. Functions returning a string result do not pass this flag, because in operations like CONCAT(a,b) we still need to evaluate precisely the collation of the result - we cannot just choose a collation of one of the sides (even if they are compatible on the given repertoire). This also applies to ExtractValue() and UpdateXML(), which now aggregate their arguments without this flag. - As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits rather than a single bit, the way to detect "is only ASCII" repertoires has changed in the code. For example: // repertoire *IS* ascii if (repertoire == MY_REPERTOIRE_ASCII) has changed in multiple places in the code to // repertoire *HAS* only ascii characters if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII)) The new function my_repertoire_is_subset_of() in m_ctype.h is used for this purpose in C code. In C++ code DTCollation::repertoire_is_subset_of() is used, e.g.: if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII)) - Repertoire of some expressions was adjusted to the new meaning: * Item_null now has MY_REPERTOIRE_NONE (was ASCII). * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30 (was EXTENDED). * Lex_string_with_metadata_st::repertoire(cs) now scans the string contents to detect the actual repertoire. * HEX() now has MY_REPERTOIRE_ASCII_ALNUM. * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale which is actually used: the explicit third argument, or @@lc_time_names. It was always @@lc_time_names before. A non-constant locale argument is assumed to be non-ASCII. * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY, JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra characters they put into the result (quotes, separators, padding, brackets). * LOWER() and UPPER() add the letters of both cases to the repertoire, because the result can contain letters of the opposite case. In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an ASCII letter can be converted to a non-ASCII one (I -> dotless i). Unicode collations are detected by the new member casefold_info_st::can_convert_to_non_ascii_on_casefolding, simple 8bit collations by to_lower['I'] and to_upper['i']. - Lex_string_with_metadata_st::repertoire(cs) now scans a string literal for any byte oriented character set (mbminlen==1), not only for ASCII based ones. It was assumed to have the full repertoire for swe7, which is not ASCII based. The scan is done by Unicode code points, so swe7 strings with digits and letters get a narrow repertoire, while e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII. Therefore swe7_swedish_ci can be mixed with other collations now, in the same way as other collations with MY_CS_IDENT_CASEUP_CI. The test func_time was adjusted: a swe7 literal ' ull' with ASCII characters is now safe to convert, the error is still tested for '[ull'. - cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS, as it sorts the minus and the dot differently. Both were verified with ctype_digits_verify. - The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL), so my_charset_is_ascii_based() is true for tis620. It was false for tis620, although tis620 is ASCII compatible. Therefore tis620 string literals were always considered to have a non-ASCII repertoire, and the repertoire optimizations did not work for them. Now ASCII-only tis620 literals are detected as ASCII. This changes the result of subselect_extra_no_semijoin from an error to success. - tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case and then compare by the code point on the entire ASCII range (the Thai specific rules affect only non-ASCII characters). A new function my_tailoring_ascii_casedn_ci() returns tailorings for such collations on ALNUM, IDENT and ASCII. For example, the underscore sorts before the letters, unlike in the caseup tailorings. - New flags were added for CHARSET_INFO::state * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the ASCII repertoire, and converts lower case letters to upper case ones, so the underscore sorts after the letters (with an upper to lower case conversion it would sort before the letters). * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations. It means that this collation has no irregularities on the IDENT repertoire (but can have irregularities say on punctuation). * MY_CS_ASCII_STD_UCA - for UCA collations. It means that a UCA collation does not reorder ASCII characters. * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits 0..9 are sorted in the code point order and are not tailored, and, for PAD collations, the space sorts before the digits. Strings of digits are compared in the same way in all such collations, even if they have irregularities on letters (latin7, cp866, czech, danish, turkish UCA, etc.), or differ by case sensitivity. It is detected from sort_order for 8bit collations, and from the tailoring rules for UCA collations. Collations with MY_CS_DIGITS_STD still must have the same PAD/NOPAD attribute. * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means MY_CS_DIGITS_STD, and additionally the minus sorts before the dot, which sorts before the digits (and the space sorts before the minus in PAD collations). Strings like '-10.5' are compared in the same way in all such collations. It is detected by my_8bit_collation_digit_flags_from_data() from sort_order for 8bit and multi-byte collations (e.g. big5_chinese_ci), and from the tailoring rules for UCA collations. The flags MY_CS_DIGITS_STD and MY_CS_ASCII_MINUS_DOT_DIGITS are not stored in ctype-extra.c, they are detected on collation initialization. - strings/conf_to_src.c was modified to detect and print MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags. - strings/ctype-extra.c was regenerated with new flags. - 8bit collations are not given MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort before the digits. A shorter string is padded with spaces for comparison, so the position of the space matters even for strings without spaces ("ab" is compared to "abc" as "ab " to "abc"). A test collation latin1_space_after_letters_ci was added to mysql-test/std_data/ldml/ for this. ctype-extra.c does not change. - The initializer of my_collation_cs_handler in ctype-utf8.c (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match my_collation_handler_st: the missing get_id and get_collation_name members were added, so that eq_collation was not initialized in their place, and the new member "tailoring" is now set to my_tailoring_none. - Results of the existing tests func_str, func_test, ps, view and subselect_sj changed from "Illegal mix of collations" to success. In these tests ASCII-only strings were compared using collations with equal tailoring on ASCII, e.g. latin1_swedish_ci with latin2_general_ci (func_str, ps), koi8r_general_ci with latin1_swedish_ci (func_test), latin1_general_ci with latin1_swedish_ci on letters and digits (view), cp932_japanese_ci with latin1_swedish_ci (subselect_sj). Such comparisons are now allowed. Cases with collations that are still incompatible on ASCII, e.g. latin7_general_ci, were added to the tests and still fail. - Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/. They display a tailoring by collation name and repertoire as returned by: cs->coll->tailoring(cs, some_repertoire) A dynamically linked plugin function collation_tailoring() was added for the purpose of these tests. Tailorings do not depend on PAD/NOPAD and on the UCA version, so the function shows "[nopad]" for NOPAD collations, and "[version X.Y.Z]" for UCA collations on the ASCII repertoire. The version is formatted by a new class UCAVersion, which uses a new method CharBuffer::append_uint8(). - Adding the test plugin/func_test/mysql-test/func_test/ ctype_caseup_ci_verify.test. It checks the order of ASCII characters for all collations declared "caseup CI" on the ASCII or IDENT range. - Adding the test plugin/func_test/mysql-test/func_test/ collation_tailoring_args.test. It checks the arguments of the plugin function collation_tailoring(): the repertoire passed by name or by number, and the NULL results for unknown names. - Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test. It displays the tailorings of collations loaded at runtime from Index.xml: simple 8bit collations (flags detected from the weights) and Unicode collations (LDML rules which do or do not reorder ASCII). - Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test They display two-dimensional charts showing which collations are compatible on which repertoires. Tests for DATE_FORMAT() with an explicit locale were added to ctype_big5_tailoring.test, ctype_cp932_tailoring.test and ctype_latin5_tailoring.test, which also has the chart for latin5. - Adding a number of MTR tests in the form of the originally reported script for various collations: SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...'; The repeated blocks of queries are shared through mysql-test/include/ ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc. - Adding tests to ctype_cp932.test checking that ExtractValue() and UpdateXML() still raise "Illegal mix of collations". |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
GoldenEmperor1177
xurbanconcept1@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-31225 Crash converting IN with a nested ROW into IN subquery What was wrong: With in_predicate_conversion_threshold set low enough, an IN predicate with a nested ROW, e.g. ROW(a,(a,a)) IN ((1,(1,1)),(2,(2,NULL))) crashed the server. Before an IN list is converted into an IN subquery over a table value constructor, cmp_row_types() checks every column with subquery_type_allows_materialization(). With a nested ROW the column is a ROW itself, and for a ROW that function must not be called. How it is fixed: A table value constructor cannot have a ROW as a column, so such a predicate cannot be converted anyway. cmp_row_types() now reports the columns as not comparable when either of them is a ROW, and the IN predicate is evaluated without the conversion. How it is tested: Added a test to main.opt_tvc with IN and NOT IN over nested ROWs. It crashes the server without the fix. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Aleksey Midenkov
midenok@gmail.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-25529 ALTER TABLE FORCE syntax improved Improves ALTER TABLE syntax when alter_list can be supplied alongside a partitioning expression, so that they can appear in any order. This is particularly useful for the FORCE clause when adding it to an existing command. Also improves handling of AUTO with FORCE, so that AUTO FORCE specified together provides more consistent syntax, which is used by this task in further commits. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Yuchen Pei
ycp@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-41064 Use my_safe_alloca for unescaped json server option string This fixes segv when the option string is longer than thread stack |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Sergei Petrunia
sergey@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
MDEV-25848: record/replay hlindex records_in_range() in optimizer context check_quick_select_array() now goes through hook_hl_records_in_range(), which records or replaces both the row estimate and the cost (record_hl_records_in_range / infuse_hl_records_in_range). The calls are stored in a new per-table "hl_records_in_range_calls" array. Co-Authored-By: Claude Sonnet 5.5 <[email protected]> |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Marko Mäkelä
marko.makela@mariadb.com |
|
|
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
fixup! fbf02b2bd3ffaff146602c102a0d5badd6e5b5b Fix a race condition in backup.backup_innodb,release |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||