Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Sutou Kouhei
MDEV-41285 ASAN heap-buffer-overflow mrn_get_string_between_quote/mrn_parse_table_param

Fix a heap buffer overflow on escaped string in table/index comment parameter (#1249)

`mrn_get_string_between_quote()` didn't advance `current_ptr` after it
processed an escape sequence such as `\x`. So it wrote the escaped
character repeatedly and overflowed the allocated buffer. It also didn't
terminate the extracted string with `\0` when the string had an escape
sequence.

For example, the following SQL caused a heap buffer overflow:

```sql
CREATE TABLE t1 (c INT) ENGINE=Mroonga COMMENT='engine "InnoDB\\x"';
```

This also returns `NULL` when `mrn_my_malloc()` fails.

Reported by Alice Sherepa. Found by Yuelin Wang.

Assisted-by: Claude:claude-5.5-opus
Oleksandr Byelkin
MDEV-40303 SIGSEGV in Item_field::type_handler() on PS re-execution

On PS re-execution a derived column is resolved via item->cached_table.
If fixing its scalar UNION subquery fails, the unit's prepare() error path
calls cleanup() but leaves 'prepared' set. find_field_in_view() returned
0 ("not found"), so find_field_in_tables() fell through to the generic
lookup, re-fixed the same subquery, prepare() returned early and
set_row() hit NULL fields.

Fix: return view_field_error when a view/derived field cannot be fixed
and stop the lookup on it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Oleksandr Byelkin
MDEV-41099 HANDLER READ missing column privilege check

HANDLER OPEN accepted column-level SELECT grants as table access, but
HANDLER READ returns all columns and never checked column privileges.
Reopen of a handler table (after FLUSH/ALTER) did no checks at all.

Require SELECT on every column (as for SELECT *) when the user has no
table-level SELECT, and redo the privilege checks on reopen.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Marko Mäkelä
fixup! 33b746888e75892654df7ffa3a696512656a631d

Calculate the correct file offset when zeroing the unused tail.
Georgi (Joro) Kodinov
more doxygen formatting added.
Andrei Elkin
MDEV-41214 PURGE BINARY LOGS does not delete non-active binlog files

PURGE BINARY LOGS could end with a warning instead of deleting files, as
the test below shows when run on the base of this commit. The warning
names a wrong reason. It follows an automatic purge that was refused by
the slave guard, slave_connections_needed_for_purge: the automatic purge
does not delete a binlog while it might still be needed by a slave, not
necessarily a connected one, that is while fewer than the variable's
number of slaves are connected.
Since MDEV-34504 the manual PURGE is meant to ignore the guard.
Yet it still protects a binlog that a slave is reading.

Technically this was possible because the automatic and the manual purge
share a code path, can_purge_log(), which cached the refusal and replayed
it, under the reason of the first arm below, to any later purge:

  if (is_active(name) ||
      (!is_relay_log && waiting_for_slave_to_change_binlog &&
      purge_sending_new_binlog_file == sending_new_binlog_file &&
      !strcmp(name, purge_binlog_name)))
    reason= "it is the current active binlog";

So "active" had to be read as "active or busy". The fix removes this
duality too: the two arms are now separate and have their own reasons.

The main issue is fixed by narrowing the busy arm with an additional
`!interactive &&` conjunct, so that a manual PURGE never consults the
cached refusal.

Running the test below on an unfixed 11.4 server (source tree at commit
42038e1145d) fails with "Result length mismatch" and, in part 2, this
diff against the recorded result:

  PURGE BINARY LOGS TO 'master-bin.000006';
  +Warnings:
  +Note 1375 Binary log 'master-bin.000003' is not purged because it is the current active binlog
  SHOW WARNINGS;
  Level Code Message
  +Note 1375 Binary log 'master-bin.000003' is not purged because it is the current active binlog
  show binary logs;
  Log_name File_size
  +master-bin.000003 #
  +master-bin.000004 #
  +master-bin.000005 #
  master-bin.000006 #
  master-bin.000007 #
  master-bin.000008 #

Test binlog.binlog_purge_stale_refusal_cache: the test sets
slave_connections_needed_for_purge=1, binlog_expire_logs_seconds=0 and
max_binlog_total_size=0 first and restores them at the end. It starts
with RESET MASTER so that the binlog names, 000003 and so on, do not depend
on the tests run before it in the same mtr environment. Part 1 is a control with
nothing cached; part 2 makes an automatic purge refuse 000003 (the files
must be older than binlog_expire_logs_seconds, hence --real_sleep 2 before
the rotation) and then runs PURGE BINARY LOGS TO 'master-bin.000006',
which must remove 000003..000005.

As a workaround, a PURGE could work if the following settings are made
before it, in this order:

  SET GLOBAL binlog_expire_logs_seconds=0, slave_connections_needed_for_purge=0;

The second setting invalidates the cached refusal. The first one keeps
the purge that this setting triggers from refusing and caching again.
Note that both settings change the server's policy, not only this one
PURGE: the time-based expiry and the slave guard are off until they are
restored.
Monty
MDEV-25292 Atomic CREATE OR REPLACE TABLE

Atomic CREATE OR REPLACE allows to keep an old table intact if the
command fails or during the crash. That is done by renaming the
original table to temporary name, as a backup and restoring it if the
CREATE fails. When the command is complete and logged the backup
table is deleted.

Atomic replace algorithm

Two DDL chains are used for CREATE OR REPLACE:
ddl_log_state_create (C) and ddl_log_state_rm (D).

  1. (C) Log rename of ORIG to TMP table (Rename TMP to original).
  2. Rename orignal to TMP.
  3. (C) Log CREATE_TABLE_ACTION of ORIG (drops ORIG);
  4. Do everything with ORIG (like insert data)
  5. (D) Log drop of TMP
  6. Write query to binlog (this marks (C) to be closed in
    case of failure)
  7. Execute drop of TMP through (D)
  8. Close (C) and (D)

  If there is a failure before 6) we revert the changes in (C)
  Chain (D) is only executed if 6) succeded (C is closed on
  crash recovery).

Foreign key errors will be found at the 1) stage.

Additional notes

- CREATE TABLE without REPLACE and temporary tables is not affected
  by this commit.
  set @@drop_before_create_or_replace=1 can be used to
  get old behaviour where existing tables are dropped
  in CREATE OR REPLACE.

- CREATE TABLE is reverted if binlogging the query fails.

- Engines having HTON_EXPENSIVE_RENAME flag set are not affected by
  this commit. Conflicting tables marked with this flag will be
  deleted with CREATE OR REPLACE.

- Replication execution is not affected by this commit.
  - Replication will first drop the conflicting table and then
    creating the new one.

- CREATE TABLE .. SELECT XID usage is fixed and now there is no need
  to log DROP TABLE via DDL_CREATE_TABLE_PHASE_LOG (see comments in
  do_postlock()). XID is now correctly updated so it disables
  DDL_LOG_DROP_TABLE_ACTION. Note that binary log is flushed at the
  final stage when the table is ready. So if we have XID in the
  binary log we don't need to drop the table.

- Three variations of CREATE OR REPLACE handled:

  1. CREATE OR REPLACE TABLE t1 (..);
  2. CREATE OR REPLACE TABLE t1 LIKE t2;
  3. CREATE OR REPLACE TABLE t1 SELECT ..;

- Test case uses 6 combinations for engines (aria, aria_notrans,
  myisam, ib, lock_tables, expensive_rename) and 2 combinations for
  binlog types (row, stmt). Combinations help to check differences
  between the results. Error failures are tested for the above three
  variations.

- expensive_rename tests CREATE OR REPLACE without atomic
  replace. The effect should be the same as with the old behaviour
  before this commit.

- Triggers mechanism is unaffected by this change. This is tested in
  create_replace.test.

- LOCK TABLES is affected. Lock restoration must be done after new
  table is created or TMP is renamed back to ORIG

- Moved ddl_log_complete() from send_eof() to finalize_ddl(). This
  checkpoint was not executed before for normal CREATE TABLE but is
  executed now.

- CREATE TABLE will now rollback also if writing to the binary
  logging failed. See rpl_gtid_strict.test

backup ddl log changes

- In case of a successfull CREATE OR REPLACE we only log
  the CREATE event, not the DROP TABLE event of the old table.

ddl_log.cc changes

  ddl_log_execute_action() now properly return error conditions.
  ddl_log_disable_entry() added to allow one to disable one entry.
  The entry on disk is still reserved until ddl_log_complete() is
  executed.

On XID usage

  Like with all other atomic DDL operations XID is used to avoid
  inconsistency between master and slave in the case of a crash after
  binary log is written and before ddl_log_state_create is closed. On
  recovery XIDs are taken from binary log and corresponding DDL log
  events get disabled.  That is done by
  ddl_log_close_binlogged_events().

On linking two chains together

  Chains are executed in the ascending order of entry_pos of execute
  entries. But entry_pos assignment order is undefined: it may assign
  bigger number for the first chain and then smaller number for the
  second chain. So the execution order in that case will be reverse:
  second chain will be executed first.

  To avoid that we link one chain to another. While the base chain
  (ddl_log_state_create) is active the secondary chain
  (ddl_log_state_rm) is not executed. That is: only one chain can be
  executed in two linked chains.

  The interface ddl_log_link_chains() was defined in "MDEV-22166
  ddl_log_write_execute_entry() extension".

Atomic info parameters in HA_CREATE_INFO

  Many functions in CREATE TABLE pass the same parameters. These
  parameters are part of table creation info and should be in
  HA_CREATE_INFO (or whatever). Passing parameters via single
  structure is much easier for adding new data and
  refactoring.

Aria changes:
- Fixed issue in Aria engine with CREATE + locked tables
  that data was not properly commited in some cases in
  case of crashes.

InnoDB related changes (by Marko Mäkelä):
- table_name_t::is_create_or_replace(): A new predicate to check for
  CREATE OR REPLACE TABLE will rename an old table to
  and eventually drop after creating the replacement.
- dict_table_t::parse_name(): Do acquire MDL on #sql-create- names
  for partitioned tables.
- dict_table_rename_in_cache(): On CREATE OR REPLACE TABLE ... SELECT,
  forget the original dict_table_t::mdl_name so that purge will
  acquire MDL on the #sql-create- name instead. In this way, the
  MDL_EXCLUSIVE that the CREATE OR REPLACE TABLE holds on the
  user-visible name will not unnecessarily block any purge of old history
  until the very end when the #sql-create- table will be dropped.
- ha_innobase::delete_table(): Do not check FOREIGN KEY consistency
  when dropping an #sql-create- table.
- row_rename_table_for_mysql(): Update SYS_FOREIGN.ID also
  when renaming to #sql-create- in order to avoid any
  duplicate key error when CREATE OR REPLACE TABLE is
  creating some FOREIGN KEY constraints by names
  that existed in the old table.

Other changes:
- Removed some auto variables in log.cc for better code readability.
- Fixed old bug that CREATE ... SELECT would not be able to auto repair
  a table that is part of the SELECT.
- Marked MyISAM that it does not support ROLLBACK (not required but
  done for better consistency with other engines).
- maria_create_trn_for_mysql() does not register a new transaction
  handler for commits. This was needed to ensure create or replace
  will not end with an active transaction.
- We do not get anymore warnings about "Engine not supporting atomic
  create" when doing a legal CREATE OR REPLACE on a table with
  foreign key constraints.
- Updated VIDEX engine flags to disable CREATE SEQUENCE.
- Removed mysql_mutex_unlock(&LOCK_gdl) / mysql_mutex_lock(&LOCK_gdl)
  around calls to binlog as these are unsafe. The binlog code uses
  global variables that needs protection from other caller.
- EITS data is preserved if create or replace fails if
  drop_before_create_or_replace=OFF. If ON, then create or replace
  will drop EITS before the drop of the original table (as before).
- Using CREATE OR REPLACE on a encrypted table that the user cannot
  decrypt will fail instead of replacing the encrypted table.
  The encrypted table will unchanged.

Known issues:
- One cannot use create or replace on an InnoDB tables that has foreign
  key references point to it
- CREATE OR REPLACE TEMPORARY table is not full atomic. Any conflicting
  table will always be dropped before creating a new one. (Old behaviour).

Bug fixes related to this MDEV:
MDEV-36435 Assertion failure in finalize_locked_tables()
MDEV-36439 Assertion `thd_arg->lex->sql_command != SQLCOM_CREATE_SEQUENCE...
MDEV-36498 Failed CoR in non-atomic mode no longer generates DROP in RBR...
MDEV-36508 Temporary files #sql-create-....frm occasionally stay after
          crash recovery
MDEV-38479 Crash in CREATE OR REPLACE SEQUENCE when new sequence cannot
          be created
MDEV-36497 Assertion failure after atomic CoR with Aria under lock in
          transactional context
MDEV-36501 EITS data is lost after failed attempt to CREATE OR REPLACE
          table
MDEV-36493 Atomic CREATE OR REPLACE ... SELECT blocks InnoDB purge
MDEV-39367 MSAN/valgrind errors in temp_file_size_cb_func,
          main.tmp_space_usage fails
MDEV-39446 Atomic CREATE OR REPLACE fails if a table cannot be decrypted
MDEV-40776 Atomic CREATE OR REPLACE silently breaks the foreign key
MDEV-40765 Assertion `"unexpected references" == 0' failed upon failing
          CREATE OR REPLACE
MDEV-40994 Atomic CoR: Transaction not registered for MariaDB 2PC, but
          transaction is active

Reverted commits:
- MDEV-36685 "CREATE-SELECT may lose in binlog side-effects of
  stored-routine" as it did not take into account that it safe to
  clear binlogs if the created table is non transactional and there
  are no other non transactional tables used.
  This was done because it caused extra logging when it is not
  needed (not using any non transactional tables) and it also did
  not solve side effects when using statement based loggging.
Sergei Golubchik
MDEV-41253 RPM %pre scriptlet unconditionally resets a pre-existing mysql user's home directory to /nonexistent

keep resetting to not writable, but use an existing path
Monty
Added --debug-dbug option to mysqltest.cc

This was to get rid of warnings when using mtr --debug
Sutou Kouhei
Fix auto directory creation for absolute mroonga_database_path_prefix (#1166)

mkdir_p() tried stat("") and mkdir("") for the leading directory
separator of an absolute path and gave up. So we couldn't create
database directory automatically for absolute
mroonga_database_path_prefix such as "/var/lib/mroonga/".

This also treats EEXIST from mkdir() as success because another process
may create the directory after our stat().

Assisted-by: Claude:claude-5.5-opus
Oleksandr Byelkin
MDEV-40303 SIGSEGV in Item_field::type_handler() on PS re-execution

On PS re-execution a derived column is resolved via item->cached_table.
If fixing its scalar UNION subquery fails, the unit's prepare() error path
calls cleanup() but leaves 'prepared' set. find_field_in_view() returned
0 ("not found"), so find_field_in_tables() fell through to the generic
lookup, re-fixed the same subquery, prepare() returned early and
set_row() hit NULL fields.

Fix:
- find_field_in_view()/find_field_in_natural_join() return the new
  field_fix_error when the found field cannot be fixed, and name
  resolution stops on it.
- st_select_lex_unit::prepare() fails for a unit whose preparation
  already failed in this execution instead of reporting it prepared.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_DIGITS - [0..9]
  * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'.
    It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and
    MY_REPERTOIRE_ASCII_DOT. They were a part of
    MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values
    of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed
    (223 and 255).
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit
  character sets no longer stops at the first bad byte sequence with
  the repertoire found so far, which could be MY_REPERTOIRE_NONE
  (compatible with everything). It returns MY_REPERTOIRE_EXTENDED
  for such strings. A unit test was added into unittest/strings.

- Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set
  can store all ASCII characters (it is false for swe7).
  It is used in Item_func_conv_charset to calculate "safe", and in
  left_is_algorithmically_simpler() to prefer other character sets
  to those where ASCII strings cannot be converted safely.

- DBUG_ASSERTs were added into mysys/charset.c, to check that
  the "tailoring" method is set in collation handlers when a collation
  is added: compiled-in, from ctype-extra.c, or loaded from Index.xml.

- Adding a new function my_collations_equal_on_repertoire(cs1, cs2,
  repertoire) in strings/. It tells if two collations are equal on
  the repertoire, taking tailoring() results and PAD/NOPAD and the UCA
  version into account. Callers do not compare tailoring() results
  directly, so the way tailorings are compared (currently by the pointer
  to the shared constant string) is private to strings/ and can be
  extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).

- DTCollation::aggregate() now makes the repertoire of the result cover
  the repertoires of both sides in all successful branches.
  Before, set(dt) replaced the repertoire with the repertoire of the
  winner side only, so the result could have a too narrow repertoire,
  e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the
  value '1.5' has a punctuation character. With the more granular
  repertoires this could wrongly allow mixing collations which are
  equal on ALNUM but different on punctuation.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale
    which is actually used: the explicit third argument, or
    @@lc_time_names. It was always @@lc_time_names before.
    A non-constant locale argument is assumed to be non-ASCII.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- Lex_string_with_metadata_st::repertoire(cs) now scans a string literal
  for any byte oriented character set (mbminlen==1), not only for
  ASCII based ones. It was assumed to have the full repertoire for swe7,
  which is not ASCII based. The scan is done by Unicode code points, so
  swe7 strings with digits and letters get a narrow repertoire, while
  e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII.
  Therefore swe7_swedish_ci can be mixed with other collations now,
  in the same way as other collations with MY_CS_IDENT_CASEUP_CI.
  The test func_time was adjusted: a swe7 literal ' ull' with ASCII
  characters is now safe to convert, the error is still tested for '[ull'.

- cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets
  the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS,
  as it sorts the minus and the dot differently. Both were verified
  with ctype_digits_verify.

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  It was false for tis620, although tis620 is ASCII compatible.
  Therefore tis620 string literals were always considered to have
  a non-ASCII repertoire, and the repertoire optimizations did not work
  for them. Now ASCII-only tis620 literals are detected as ASCII.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case
  and then compare by the code point on the entire ASCII range
  (the Thai specific rules affect only non-ASCII characters).
  A new function my_tailoring_ascii_casedn_ci() returns tailorings
  for such collations on ALNUM, IDENT and ASCII. For example, the
  underscore sorts before the letters, unlike in the caseup tailorings.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire, and converts lower case letters to upper case ones,
    so the underscore sorts after the letters (with an upper to lower
    case conversion it would sort before the letters).

  * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

  * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits
    0..9 are sorted in the code point order and are not tailored, and,
    for PAD collations, the space sorts before the digits. Strings of
    digits are compared in the same way in all such collations, even if
    they have irregularities on letters (latin7, cp866, czech, danish,
    turkish UCA, etc.), or differ by case sensitivity. It is detected
    from sort_order for 8bit collations, and from the tailoring rules
    for UCA collations. Collations with MY_CS_DIGITS_STD still must have
    the same PAD/NOPAD attribute.

  * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means
    MY_CS_DIGITS_STD, and additionally the minus sorts before the dot,
    which sorts before the digits (and the space sorts before the minus
    in PAD collations). Strings like '-10.5' are compared in the same
    way in all such collations. It is detected from sort_order for simple
    8bit collations (my_coll_init_simple()), and from the tailoring rules
    for UCA collations. For the multi-byte collations big5_chinese_ci,
    gbk_chinese_ci, gb2312_chinese_ci and sjis_japanese_ci it is declared
    by my_tailoring_ident_caseup_ci_generic(). The flags MY_CS_DIGITS_STD
    and MY_CS_ASCII_MINUS_DOT_DIGITS are stored in ctype-extra.c for
    compiled collations (conf_to_src.c prints them), and the detection is
    skipped for a collation which already has any of the repertoire flags
    (MY_CS_REPERTOIRE_FLAGS).

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- 8bit collations are not given MY_CS_ASCII_CASEUP_CI and
  MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort
  before the digits. A shorter string is padded with spaces for
  comparison, so the position of the space matters even for strings
  without spaces ("ab" is compared to "abc" as "ab " to "abc").
  A test collation latin1_space_after_letters_ci was added to
  mysql-test/std_data/ldml/ for this. ctype-extra.c does not change.

- The initializer of my_collation_cs_handler in ctype-utf8.c
  (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match
  my_collation_handler_st: the missing get_id and get_collation_name
  members were added, so that eq_collation was not initialized in their
  place, and the new member "tailoring" is now set to my_tailoring_none.

- Results of the existing tests func_str, func_test, ps, view
  and subselect_sj changed from "Illegal mix of collations" to success.
  In these tests ASCII-only strings were compared using collations
  with equal tailoring on ASCII, e.g. latin1_swedish_ci with
  latin2_general_ci (func_str, ps), koi8r_general_ci with
  latin1_swedish_ci (func_test), latin1_general_ci with
  latin1_swedish_ci on letters and digits (view), cp932_japanese_ci
  with latin1_swedish_ci (subselect_sj).
  Such comparisons are now allowed.
  Cases with collations that are still incompatible on ASCII,
  e.g. latin7_general_ci, were added to the tests and still fail.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_caseup_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "caseup CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.
  Tests for DATE_FORMAT() with an explicit locale were added to
  ctype_big5_tailoring.test, ctype_cp932_tailoring.test and
  ctype_latin5_tailoring.test, which also has the chart for latin5.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
Sergei Golubchik
workaround for https://bugzilla.redhat.com/show_bug.cgi?id=2390105
iff-sal
MDEV-40296: SELECT @@GLOBAL.replicate_do_db ignores default_master_connection

When querying replication filter system variables without specifying a
connection name qualifier (e.g., SELECT @@GLOBAL.replicate_do_db),
Sys_var_rpl_filter::global_value_ptr() previously passed the unspecified
base_name directly to get_master_info(), which resolved to the default
connection (empty_clex_str) rather than respecting
@@SESSION.default_master_connection.

Fix by falling back to thd->variables.default_master_connection when
base_name->str is NULL in Sys_var_rpl_filter::global_value_ptr().

Add regression test in suite/multi_source/rpl_filter_default_master.test
covering explicit named connection, unspecified connection falling back
to default, and explicit unnamed connection with empty backticks.
Marko Mäkelä
fixup! 276b6dd3a08fec92e474963e8defecd22bc092ad

error C2886: 'tpool::pwrite': symbol cannot be used in a member using-declaration
Georgi (Joro) Kodinov
more doxygen formatting added.
Sergei Golubchik
MDEV-41431 LOAD_FILE checks for is_secure_file_path() one path but opens another
Vladislav Vaintroub
MDEV-39150 Some data conversion macros fail to use memcpy()

MDEV-37788 converted the uintNkorr() and intNstore() macros to use
memcpy() and byte swap intrinsics, but the floating point and short/long
conversion macros in big_endian.h and myisampack.h still accessed data
one byte at a time, and those in little_endian.h were a mix of both.
Compilers may emit slow byte-at-a-time loads and stores for such code.

Reimplement the macros with memcpy(), and with MY_BSWAP32/MY_BSWAP64
where the byte order differs from the host. big_endian.h and
little_endian.h no longer differ, so move the macros to my_byteorder.h
and remove the two headers, which are no longer installed. Implement
mi_float4store() etc. in myisampack.h with the same helper macros.

Remove the unused ulongget() macro, and the code for the mixed-endian
floating point layout (a little-endian CPU with big-endian floating
point word order) from the macros, change_double_for_sort() and dtoa.c.
It only applied to the obsolete ARM FPA format.

Reimplement mach_double_read(), mach_double_write(), mach_float_read()
and mach_float_write() in InnoDB with float8get(), float8store(),
float4get() and float4store(), instead of copying bytes in a loop.

The stored formats do not change. The unit test byte_order-t now checks
the byte layout of the floating point macros and the sign extension of
the native byte order macros, so no MTR test is added.
Kristian Nielsen
MDEV-41217: Inefficient memory management by parallel slave

Draft patch / proof-of-concept.

This patch splits the rpl_group_info into two parts. A small part that
contains only the information necessary in the queue of events sent from the
SQL driver thread to parallel replication worker threads is made separate
and pointed to by rpl_group_info::q.

This greatly reduces the memory requirements for the queued event groups,
especially when the slave is lagging and there are a lot of event groups
queued (eg. due to large --slave-parallel-max-queued).

Now the main rpl_group_info object only exists once per worker thread (and
once in the SQL driver thread).

Signed-off-by: Kristian Nielsen <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_DIGITS - [0..9]
  * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'.
    It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and
    MY_REPERTOIRE_ASCII_DOT. They were a part of
    MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values
    of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed
    (223 and 255).
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit
  character sets no longer stops at the first bad byte sequence with
  the repertoire found so far, which could be MY_REPERTOIRE_NONE
  (compatible with everything). It returns MY_REPERTOIRE_EXTENDED
  for such strings. A unit test was added into unittest/strings.

- Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set
  can store all ASCII characters (it is false for swe7).
  It is used in Item_func_conv_charset to calculate "safe", and in
  left_is_algorithmically_simpler() to prefer other character sets
  to those where ASCII strings cannot be converted safely.

- DBUG_ASSERTs were added into mysys/charset.c, to check that
  the "tailoring" method is set in collation handlers when a collation
  is added: compiled-in, from ctype-extra.c, or loaded from Index.xml.

- Adding a new function my_collations_equal_on_repertoire(cs1, cs2,
  repertoire) in strings/. It tells if two collations are equal on
  the repertoire, taking tailoring() results and PAD/NOPAD and the UCA
  version into account. Callers do not compare tailoring() results
  directly, so the way tailorings are compared (currently by the pointer
  to the shared constant string) is private to strings/ and can be
  extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).

- DTCollation::aggregate() now makes the repertoire of the result cover
  the repertoires of both sides in all successful branches.
  Before, set(dt) replaced the repertoire with the repertoire of the
  winner side only, so the result could have a too narrow repertoire,
  e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the
  value '1.5' has a punctuation character. With the more granular
  repertoires this could wrongly allow mixing collations which are
  equal on ALNUM but different on punctuation.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale
    which is actually used: the explicit third argument, or
    @@lc_time_names. It was always @@lc_time_names before.
    A non-constant locale argument is assumed to be non-ASCII.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- Lex_string_with_metadata_st::repertoire(cs) now scans a string literal
  for any byte oriented character set (mbminlen==1), not only for
  ASCII based ones. It was assumed to have the full repertoire for swe7,
  which is not ASCII based. The scan is done by Unicode code points, so
  swe7 strings with digits and letters get a narrow repertoire, while
  e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII.
  Therefore swe7_swedish_ci can be mixed with other collations now,
  in the same way as other collations with MY_CS_IDENT_CASEUP_CI.
  The test func_time was adjusted: a swe7 literal ' ull' with ASCII
  characters is now safe to convert, the error is still tested for '[ull'.

- cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets
  the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS,
  as it sorts the minus and the dot differently. Both were verified
  with ctype_digits_verify.

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  It was false for tis620, although tis620 is ASCII compatible.
  Therefore tis620 string literals were always considered to have
  a non-ASCII repertoire, and the repertoire optimizations did not work
  for them. Now ASCII-only tis620 literals are detected as ASCII.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case
  and then compare by the code point on the entire ASCII range
  (the Thai specific rules affect only non-ASCII characters).
  A new function my_tailoring_ascii_casedn_ci() returns tailorings
  for such collations on ALNUM, IDENT and ASCII. For example, the
  underscore sorts before the letters, unlike in the caseup tailorings.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire, and converts lower case letters to upper case ones,
    so the underscore sorts after the letters (with an upper to lower
    case conversion it would sort before the letters).

  * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

  * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits
    0..9 are sorted in the code point order and are not tailored, and,
    for PAD collations, the space sorts before the digits. Strings of
    digits are compared in the same way in all such collations, even if
    they have irregularities on letters (latin7, cp866, czech, danish,
    turkish UCA, etc.), or differ by case sensitivity. It is detected
    from sort_order for 8bit collations, and from the tailoring rules
    for UCA collations. Collations with MY_CS_DIGITS_STD still must have
    the same PAD/NOPAD attribute.

  * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means
    MY_CS_DIGITS_STD, and additionally the minus sorts before the dot,
    which sorts before the digits (and the space sorts before the minus
    in PAD collations). Strings like '-10.5' are compared in the same
    way in all such collations. It is detected from sort_order for simple
    8bit collations (my_coll_init_simple()), and from the tailoring rules
    for UCA collations. For the multi-byte collations big5_chinese_ci,
    gbk_chinese_ci, gb2312_chinese_ci and sjis_japanese_ci it is declared
    by my_tailoring_ident_caseup_ci_generic(). The flags MY_CS_DIGITS_STD
    and MY_CS_ASCII_MINUS_DOT_DIGITS are stored in ctype-extra.c for
    compiled collations (conf_to_src.c prints them), and the detection is
    skipped for a collation which already has any of the repertoire flags
    (MY_CS_REPERTOIRE_FLAGS).

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- 8bit collations are not given MY_CS_ASCII_CASEUP_CI and
  MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort
  before the digits. A shorter string is padded with spaces for
  comparison, so the position of the space matters even for strings
  without spaces ("ab" is compared to "abc" as "ab " to "abc").
  A test collation latin1_space_after_letters_ci was added to
  mysql-test/std_data/ldml/ for this. ctype-extra.c does not change.

- The initializer of my_collation_cs_handler in ctype-utf8.c
  (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match
  my_collation_handler_st: the missing get_id and get_collation_name
  members were added, so that eq_collation was not initialized in their
  place, and the new member "tailoring" is now set to my_tailoring_none.

- Results of the existing tests func_str, func_test, ps, view
  and subselect_sj changed from "Illegal mix of collations" to success.
  In these tests ASCII-only strings were compared using collations
  with equal tailoring on ASCII, e.g. latin1_swedish_ci with
  latin2_general_ci (func_str, ps), koi8r_general_ci with
  latin1_swedish_ci (func_test), latin1_general_ci with
  latin1_swedish_ci on letters and digits (view), cp932_japanese_ci
  with latin1_swedish_ci (subselect_sj).
  Such comparisons are now allowed.
  Cases with collations that are still incompatible on ASCII,
  e.g. latin7_general_ci, were added to the tests and still fail.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_caseup_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "caseup CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.
  Tests for DATE_FORMAT() with an explicit locale were added to
  ctype_big5_tailoring.test, ctype_cp932_tailoring.test and
  ctype_latin5_tailoring.test, which also has the chart for latin5.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
Khaled Riyad
MDEV-41140 mariadb-server.postinst chown dereferences symlinks

The postinst and preinst scripts followed symlinks when fixing
ownership under /var/log/mysql and /var/lib/mysql. Use -h and stop
find from descending so symlinks there are no longer dereferenced.
Oleksandr Byelkin
MDEV-32401 expression cache lead to crash if table of wrong type created and the cache switched off

1) take into account TMP_TABLE_ALL_COLUMNS when
  we are modifying agg_item->result_field
2) remove unused now "bool materialized_subquery;"

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Vladislav Vaintroub
MDEV-39150 Some data conversion macros fail to use memcpy()

MDEV-37788 converted the uintNkorr() and intNstore() macros to use
memcpy() and byte swap intrinsics, but the floating point and short/long
conversion macros in big_endian.h and myisampack.h still accessed data
one byte at a time, and those in little_endian.h were a mix of both.
Compilers may emit slow byte-at-a-time loads and stores for such code.

Reimplement the macros with memcpy(), and with MY_BSWAP32/MY_BSWAP64
where the byte order differs from the host. big_endian.h and
little_endian.h no longer differ, so move the macros to my_byteorder.h
and remove the two headers, which are no longer installed. Implement
mi_float4store() etc. in myisampack.h with the same helper macros.

Remove the unused ulongget() macro, and the code for the mixed-endian
floating point layout (a little-endian CPU with big-endian floating
point word order) from the macros, change_double_for_sort() and dtoa.c.
It only applied to the obsolete ARM FPA format.

Reimplement mach_double_read(), mach_double_write(), mach_float_read()
and mach_float_write() in InnoDB with float8get(), float8store(),
float4get() and float4store(), instead of copying bytes in a loop.

The stored formats do not change. The unit test byte_order-t now checks
the byte layout of the floating point macros and the sign extension of
the native byte order macros, so no MTR test is added.
Oleg Smirnov
MDEV-32412 Pushdown from HAVING: Item_func is immutable while arguments are not

Fix-up for commit e97560eac03 (MDEV-28958), which made set_extraction_flag()
ignore basic constants so that the read-only Item_true/Item_false singletons
are never written to. Because of that, the callers that mark a subtree with
MARKER_IMMUTABLE before pushing a condition from HAVING into WHERE cannot
mark basic constants inside it.

Item::cleanup_excluding_immutables_processor() did not know about this
exception and cleaned such items up, unfixing them while their marked parents
stayed fixed. fix_fields() does not descend into fixed items, so the constant
was left unfixed and Item_direct_view_ref::used_tables() later dereferenced a
NULL null_ref_table. Skip basic constants there too.
Pekka Lampio
MDEV-40018: SHOW SLAVE STATUS allocates ~1.7 MB per call due to oversized VARCHAR columns

SHOW SLAVE STATUS / SHOW ALL SLAVES STATUS / INFORMATION_SCHEMA.SLAVE_STATUS
are all served from the ST_FIELD_INFO array slave_status_info[]. Thirteen of
its columns (Replicate_Do_DB, Replicate_Ignore_DB, Replicate_Do_Table,
Replicate_Ignore_Table, Replicate_Wild_Do_Table, Replicate_Wild_Ignore_Table,
Replicate_Ignore_Server_Ids, Gtid_IO_Pos, Replicate_Do_Domain_Ids,
Replicate_Ignore_Domain_Ids, Slave_SQL_Running_State, Replicate_Rewrite_DB,
and Gtid_Slave_Pos) used the no-argument Varchar() helper, which defaults to
MAX_FIELD_VARCHARLENGTH/3 characters. With utf8mb3, that reserves 65532
bytes per column in the temporary table's record buffer regardless of the
actual value size, and the record is allocated twice, accounting for the
reported ~1.7 MB allocated on every call.

Change these columns to Blob(MAX_FIELD_VARCHARLENGTH), which keeps the same
maximum capacity but stores only a pointer and length in the record. Values,
NULL-ability, and column order are unchanged; the client-visible column type
for these columns changes from VARCHAR to BLOB.

mysql-test/suite/funcs_1/r/is_columns_is.result is updated to match, since
it asserts on INFORMATION_SCHEMA.COLUMNS metadata for these columns.

Assisted-By: Claude AI
Sergei Golubchik
MDEV-39624 storage/connect/CMakeLists.txt: unguarded generator expression variable causes fatal error at cmake generate phase
Sergei Golubchik
MDEV-39625 plugin/auth_pam/testing/CMakeLists.txt breaks INSTALL_MYSQLTESTDIR= suppression
Kristian Nielsen
Small docs update for binlog-in-engine

Signed-off-by: Kristian Nielsen <[email protected]>
Monty
ha_table_exists() cleanup and improvement

This is part of MDEV-25292 Atomic CREATE OR REPLACE TABLE.

Removed default values for arguments, added flags argument to specify
filename flags (FN_TO_IS_TMP, FN_FROM_IS_TMP) and forward the flag to
build_table_name().

Original patch from: Aleksey Midenkov <[email protected]>
Monty
MDEV-31138 Enable the test spider/bugfix.mdev_29676
Oleksandr Byelkin
MDEV-41099 HANDLER READ missing column privilege check

HANDLER OPEN accepted column-level SELECT grants as table access, but
HANDLER READ returns all columns and never checked column privileges.
Reopen of a handler table (after FLUSH/ALTER) did no checks at all.

Require SELECT on every column (as for SELECT *) when the user has no
table-level SELECT, and redo the privilege checks on reopen.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Monty
MDEV-23298 Assertion `table_list->prelocking_placeholder == TABLE_LIST::PRELOCK_NONE' failed in check_lock_and_start_stmt on CREATE OR REPLACE TABLE

Fixed by removing wrong assert

Review: Sanja Byelkin
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_DIGITS - [0..9]
  * MY_REPERTOIRE_ASCII_MINUS_DOT_DIGITS - [-.0..9], e.g. '-10.5'.
    It consists of the new bits MY_REPERTOIRE_ASCII_MINUS and
    MY_REPERTOIRE_ASCII_DOT. They were a part of
    MY_REPERTOIRE_ASCII_NOT_IDENT before. As a result, the numeric values
    of MY_REPERTOIRE_ASCII and MY_REPERTOIRE_UNICODE30 changed
    (223 and 255).
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- my_string_repertoire() for ucs2, utf16, utf32 and NONASCII 8bit
  character sets no longer stops at the first bad byte sequence with
  the repertoire found so far, which could be MY_REPERTOIRE_NONE
  (compatible with everything). It returns MY_REPERTOIRE_EXTENDED
  for such strings. A unit test was added into unittest/strings.

- Adding CHARSET_INFO::is_ascii_superset(). It tells if a character set
  can store all ASCII characters (it is false for swe7).
  It is used in Item_func_conv_charset to calculate "safe", and in
  left_is_algorithmically_simpler() to prefer other character sets
  to those where ASCII strings cannot be converted safely.

- DBUG_ASSERTs were added into mysys/charset.c, to check that
  the "tailoring" method is set in collation handlers when a collation
  is added: compiled-in, from ctype-extra.c, or loaded from Index.xml.

- Adding a new function my_collations_equal_on_repertoire(cs1, cs2,
  repertoire) in strings/. It tells if two collations are equal on
  the repertoire, taking tailoring() results and PAD/NOPAD and the UCA
  version into account. Callers do not compare tailoring() results
  directly, so the way tailorings are compared (currently by the pointer
  to the shared constant string) is private to strings/ and can be
  extended later, e.g. for UCA collations on MY_REPERTOIRE_UNICODE30.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).

- DTCollation::aggregate() now makes the repertoire of the result cover
  the repertoires of both sides in all successful branches.
  Before, set(dt) replaced the repertoire with the repertoire of the
  winner side only, so the result could have a too narrow repertoire,
  e.g. IF(1, 1.5, HEX(255)) had MY_REPERTOIRE_ASCII_ALNUM although the
  value '1.5' has a punctuation character. With the more granular
  repertoires this could wrongly allow mixing collations which are
  equal on ALNUM but different on punctuation.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale
    which is actually used: the explicit third argument, or
    @@lc_time_names. It was always @@lc_time_names before.
    A non-constant locale argument is assumed to be non-ASCII.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- Lex_string_with_metadata_st::repertoire(cs) now scans a string literal
  for any byte oriented character set (mbminlen==1), not only for
  ASCII based ones. It was assumed to have the full repertoire for swe7,
  which is not ASCII based. The scan is done by Unicode code points, so
  swe7 strings with digits and letters get a narrow repertoire, while
  e.g. 0x5B (LATIN CAPITAL LETTER A WITH DIAERESIS in swe7) is non-ASCII.
  Therefore swe7_swedish_ci can be mixed with other collations now,
  in the same way as other collations with MY_CS_IDENT_CASEUP_CI.
  The test func_time was adjusted: a swe7 literal ' ull' with ASCII
  characters is now safe to convert, the error is still tested for '[ull'.

- cp1250_czech_cs sorts -, . and 0..9 in the code point order, so it gets
  the MINUS_DOT_DIGITS tailoring. latin2_czech_cs gets only DIGITS,
  as it sorts the minus and the dot differently. Both were verified
  with ctype_digits_verify.

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  It was false for tis620, although tis620 is ASCII compatible.
  Therefore tis620 string literals were always considered to have
  a non-ASCII repertoire, and the repertoire optimizations did not work
  for them. Now ASCII-only tis620 literals are detected as ASCII.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case
  and then compare by the code point on the entire ASCII range
  (the Thai specific rules affect only non-ASCII characters).
  A new function my_tailoring_ascii_casedn_ci() returns tailorings
  for such collations on ALNUM, IDENT and ASCII. For example, the
  underscore sorts before the letters, unlike in the caseup tailorings.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire, and converts lower case letters to upper case ones,
    so the underscore sorts after the letters (with an upper to lower
    case conversion it would sort before the letters).

  * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

  * MY_CS_DIGITS_STD - for any collation. It means that the ASCII digits
    0..9 are sorted in the code point order and are not tailored, and,
    for PAD collations, the space sorts before the digits. Strings of
    digits are compared in the same way in all such collations, even if
    they have irregularities on letters (latin7, cp866, czech, danish,
    turkish UCA, etc.), or differ by case sensitivity. It is detected
    from sort_order for 8bit collations, and from the tailoring rules
    for UCA collations. Collations with MY_CS_DIGITS_STD still must have
    the same PAD/NOPAD attribute.

  * MY_CS_ASCII_MINUS_DOT_DIGITS - for any collation. Means
    MY_CS_DIGITS_STD, and additionally the minus sorts before the dot,
    which sorts before the digits (and the space sorts before the minus
    in PAD collations). Strings like '-10.5' are compared in the same
    way in all such collations. It is detected by
    my_8bit_collation_digit_flags_from_data() from sort_order for 8bit
    and multi-byte collations (e.g. big5_chinese_ci), and from
    the tailoring rules for UCA collations. The flags
    MY_CS_DIGITS_STD and MY_CS_ASCII_MINUS_DOT_DIGITS are not stored
    in ctype-extra.c, they are detected on collation initialization.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- 8bit collations are not given MY_CS_ASCII_CASEUP_CI and
  MY_CS_IDENT_CASEUP_CI if they are PAD and the space does not sort
  before the digits. A shorter string is padded with spaces for
  comparison, so the position of the space matters even for strings
  without spaces ("ab" is compared to "abc" as "ab " to "abc").
  A test collation latin1_space_after_letters_ci was added to
  mysql-test/std_data/ldml/ for this. ctype-extra.c does not change.

- The initializer of my_collation_cs_handler in ctype-utf8.c
  (utf8mb3_general_cs, under HAVE_UTF8_GENERAL_CS) was fixed to match
  my_collation_handler_st: the missing get_id and get_collation_name
  members were added, so that eq_collation was not initialized in their
  place, and the new member "tailoring" is now set to my_tailoring_none.

- Results of the existing tests func_str, func_test, ps, view
  and subselect_sj changed from "Illegal mix of collations" to success.
  In these tests ASCII-only strings were compared using collations
  with equal tailoring on ASCII, e.g. latin1_swedish_ci with
  latin2_general_ci (func_str, ps), koi8r_general_ci with
  latin1_swedish_ci (func_test), latin1_general_ci with
  latin1_swedish_ci on letters and digits (view), cp932_japanese_ci
  with latin1_swedish_ci (subselect_sj).
  Such comparisons are now allowed.
  Cases with collations that are still incompatible on ASCII,
  e.g. latin7_general_ci, were added to the tests and still fail.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_caseup_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "caseup CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.
  Tests for DATE_FORMAT() with an explicit locale were added to
  ctype_big5_tailoring.test, ctype_cp932_tailoring.test and
  ctype_latin5_tailoring.test, which also has the chart for latin5.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
Sergei Golubchik
MDEV-41449 ASAN server crash on corrupted frm

make sure the declared (not actual) file length is not too small
Alessandro Vetere
fixup! MDEV-33966: buf_page_make_young() is a contention point
Sutou Kouhei
MDEV-41287 ASAN: stack-buffer-overflow in strcpy/mrn::DatabaseRepairer::detect_paths/mrn::DatabaseRepairer::each_database

Fix a buffer overflow with long mroonga_database_path_prefix (#1159)

Reported by Yuelin Wang. Thanks!!!

We used fixed size buffers (MRN_MAX_PATH_SIZE) for database paths built
from mroonga_database_path_prefix. So a long
mroonga_database_path_prefix caused a buffer overflow.

This uses std::string for them. We don't need to check path length in
Mroonga because Groonga reports an error for too long path.

Assisted-by: Claude:claude-5.5-opus
GoldenEmperor1177
MDEV-31225 Crash converting IN with a nested ROW into IN subquery

What was wrong:
With in_predicate_conversion_threshold set low enough, an IN predicate
with a nested ROW, e.g.

  ROW(a,(a,a)) IN ((1,(1,1)),(2,(2,NULL)))

crashed the server. Before an IN list is converted into an IN subquery
over a table value constructor, cmp_row_types() checks every column
with subquery_type_allows_materialization(). With a nested ROW the
column is a ROW itself, and for a ROW that function must not be called.

How it is fixed:
A table value constructor cannot have a ROW as a column, so such a
predicate cannot be converted anyway. cmp_row_types() now reports the
columns as not comparable when either of them is a ROW, and the IN
predicate is evaluated without the conversion.

How it is tested:
Added a test to main.opt_tvc with IN and NOT IN over nested ROWs. It
crashes the server without the fix.
Sergei Golubchik
cleanup: put "bad frm" tests in one file
Sergei Petrunia
MDEV-25848: record/replay hlindex records_in_range() in optimizer context

check_quick_select_array() now goes through hook_hl_records_in_range(),
which records or replaces both the row estimate and the cost
(record_hl_records_in_range / infuse_hl_records_in_range). The calls are
stored in a new per-table "hl_records_in_range_calls" array.

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Marko Mäkelä
fixup! fbf02b2bd3ffaff146602c102a0d5badd6e5b5b

Fix a race condition in backup.backup_innodb,release