journal · 2 august 2026
Two merged patches
Notes on two patches merged into projects I had not worked on before. One is a grammar change in Java, the other is a test migration in Python.
Apache ShardingSphere: Doris partition syntax
- Project
- Apache ShardingSphere, a distributed SQL layer that shards, routes and rewrites queries
- Language
- Java, ANTLR4 grammar
- Change
- 3 files, 73 lines added, 2 removed
- Merged
- 30 July 2026 into master
ShardingSphere parses SQL before it can route or rewrite it. The parsing is driven by ANTLR4 grammar files, one set per SQL dialect, which describe every statement the engine accepts. If a dialect has syntax the grammar does not describe, the parser fails on that statement.
Doris is one of the supported dialects, and its grammar had a gap around partitions. An open issue listed statements that were failing to parse. Three kinds of partition syntax were falling through:
-- multi range: generate many partitions from a range and a step
PARTITION BY RANGE (dt) (
FROM ("2022-01-01") TO ("2022-01-31") INTERVAL 1 DAY
)
-- fixed range: an explicit half open interval
PARTITION p1 VALUES [("2022-01-01"), ("2022-02-01"))
-- per partition properties after the values clause
PARTITION p1 VALUES LESS THAN ("2022-01-01")
("storage_policy" = "cooldown", "replication_num" = "1")
None of this was new syntax to the project. Doris already accepted all three forms in ALTER TABLE, and the grammar already had rules for them. They had not been wired into the CREATE TABLE path. So the change mostly reuses rules that already existed.
The main part was splitting the list of partition definitions so an entry could be either an ordinary partition or a multi range expansion:
// before
partitionDefinitions
: LP_ partitionDefinition (COMMA_ partitionDefinition)* RP_
;
// after
partitionDefinitions
: LP_ partitionDefinitionItem (COMMA_ partitionDefinitionItem)* RP_
;
partitionDefinitionItem
: partitionDefinition | dorisMultiRangePartition
;
dorisMultiRangePartition
: FROM LP_ expr RP_ TO LP_ expr RP_ INTERVAL expr intervalUnit?
;
Then partitionDefinition itself gained the fixed range alternative inside its VALUES clause and an optional (LP_ properties RP_) for the per partition property list, reusing the existing properties rule.
ShardingSphere marks dialect specific edits inside shared grammar with // DORIS CHANGED BEGIN and // DORIS ADDED BEGIN comment fences. This is not in the contributing guide. I found it by reading the surrounding file. A first revision that used the wrong fence was flagged in review and corrected.
The rest of the diff, 56 of the 73 lines, is test cases: SQL samples plus the expected parse assertions in the project's XML fixture format. A grammar change needs them, because the claim being made is that a statement which used to fail now parses into a specific tree.
Before opening the PR I ran the Doris parser integration suite locally and put the result in the description: 1261 tests, zero failures. The issue also listed a few SQL fragments that were not valid standalone statements, plus one case that already parsed on master. I said so in the description rather than skipping them silently, so it was clear why there were four assertions against a list of seven cases.
Lightly: unittest to pytest
- Project
- Lightly, a Python library for self supervised learning on images
- Language
- Python, pytest
- Change
- 7 files, 248 lines added, 301 removed
- Merged
- 28 July 2026 into master
This came from an umbrella issue. The maintainers wanted the test suite moved from unittest.TestCase to plain pytest, split into groups so several people could work in parallel without colliding. I took the models group, seven files.
Most of it is mechanical. Drop the TestCase base class, turn setUp into an autouse fixture, replace self.assertEqual with a bare assert, replace assertRaises with pytest.raises, and swap @unittest.skipUnless for @pytest.mark.skipif.
The largest change was to how the suite handled GPU tests. Every device sensitive test existed twice: a real test defaulting to CPU, and a CUDA wrapper that called it back with different arguments.
# before: two methods, one of them a shim
def test_single_projection_head(self, device: str = "cpu", seed: int = 0) -> None:
...
@unittest.skipUnless(torch.cuda.is_available(), "skip")
def test_single_projection_head_cuda(self, seed: int = 0) -> None:
self.test_single_projection_head(device="cuda", seed=seed)
# after: one method, two parametrized runs
@pytest.mark.parametrize("device", ["cpu", "cuda"])
def test_single_projection_head(self, device: str) -> None:
if device == "cuda" and not torch.cuda.is_available():
pytest.skip("CUDA not available")
...
That is where most of the 301 deleted lines went. The nested subTest loops came out as well, since pytest reports a failing parametrized case on its own.
The point of a refactor PR is that nothing changes except the structure. So when a test looked wrong I left it, and when a name was poor I kept it. Any improvement made along the way would be something the reviewer had to check separately, on top of a diff that is otherwise mechanical.
There was also an existing config detail to respect: five of the seven files are excluded from mypy, and two are not. I left the exclusions as they were, matched what the sibling PRs in the same umbrella issue had done, and confirmed in the description that ruff, mypy and pytest were all clean on the directory.
Notes
- Both patches had a precedent in the repo already. ShardingSphere had the same partition rules working under ALTER TABLE, and Lightly had sibling migration PRs plus a reference file the maintainers pointed to.
- Conventions like the comment fences and the mypy exclusions are not what I would have chosen, but they are what the projects use.
- I ran the test suites before opening each PR and put the counts in the description.
- Both descriptions explain something I chose not to do. That saves a round trip where a reviewer asks why a case is missing.
- Both diffs are small. In a codebase I do not know, that is easier for a reviewer to check.