public class PageTreeFactory {
public static Page loadPageTree(/*...*/) {
// [...]
}
}| Data-Oriented Programming |
| A Lengthy Example |
| Project Amber's Arc |
Slides at slides.nipafx.dev.
| Data-Oriented Programming |
| A Lengthy Example |
| Project Amber's Arc |
Data-oriented programming:
focuses on data (Duh!)
views programs as series of transformations
Consider it for:
smaller (sub)systems
that mainly process input to output
Brian Goetz formulated them in June 2022:
Data Oriented Programming in Java
I offered an updated version in May 2024:
Data-Oriented Programming in Java - Version 1.1
Version 1.1:
Model data immutably and transparently.
Model the data, the whole data,
and nothing but the data.
Make illegal states unrepresentable.
Separate operations from data.
| Data-Oriented Programming |
| A Lengthy Example |
| Project Amber's Arc |
Starting with a seed URL:
connect to URL
identify kind of page
identify interesting section
identify outgoing links
for each link, start at 1.
(Code on github.com/nipafx/modern-java-demo.)
That logic is implemented in:
public class PageTreeFactory {
public static Page loadPageTree(/*...*/) {
// [...]
}
}What does Page look like?
Pages:
all pages have a url
unresolved pages have an error
resolved pages have content
GitHub pages have:
outgoing links
issueNumber or prNumber
Operations:
evaluate statistics
create pretty string
A single Page class with this API:
public URL url();
public Exception error();
public String content();
public int issueNumber();
public int prNumber();
public Set<Page> links();
public Stats evaluateStatistics();
public String toPrettyString();Problems:
page "type" is implicit
legal combination of fields is unclear
clients must "divine" the type
disparate operations on same class
Data-oriented programming:
makes all data explicit and "obviously correct"
separates operations from data
models systems as production line
Model data immutably and transparently.
Records fit perfectly:
record ExternalPage(URI url, String content) { }Records are shallowly immutable,
but field types may not be.
⇝ Fix that during construction.
Model the data, the whole data,
and nothing but the data.
There are four kinds of pages:
error page
external page
GitHub issue page
GitHub PR page
⇝ Use four records to model them!
public record ErrorPage(
URI url, Exception ex) { }
public record ExternalPage(
URI url, String content) { }
public record GitHubIssuePage(
URI url, String content,
int issueNumber, Set<Page> links) { }
public record GitHubPrPage(
URI url, String content,
int prNumber, Set<Page> links) { }There are additional relations between them:
a page (load) is either successful or not
a successful page is either external or GitHub
a GitHub page is either for a PR or an issue
⇝ Use sealed types to model the alternatives!
public sealed interface Page
permits ErrorPage, SuccessfulPage {
URI url();
}
public sealed interface SuccessfulPage
extends Page permits ExternalPage, GitHubPage {
String content();
}
public sealed interface GitHubPage
extends SuccessfulPage
permits GitHubIssuePage, GitHubPrPage {
Set<Page> links();
}Make illegal states unrepresentable.
Many are already, e.g.:
with error and with content
with issueNumber and prNumber
with isseNumber or prNumber but no links
⇝ Reject other illegal states in constructors.
page "type" is explicit in Java’s type
only legal combination of fields are possible
API is more self-documenting
code is easier to test
But where did the operations go?
Separate operations from data.
⇝ Record methods should be limited to derived quantities.
public Stats evaluateStatistics();
public String toPrettyString();This actually applies to our operations.
But what if it didn’t? 😁
Pattern matching on sealed types is perfect
to apply polymorphic operations to data!
And records eschew encapsulation,
so everything is accessible.
In class Statistician:
private void evaluatePage(Page page) {
// `numberOf...` are fields
switch (page) {
case GitHubIssuePage issue -> numberOfIssues++;
case GitHubPrPage pr -> numberOfPrs++;
case ExternalPage ext -> numberOfExternals++;
case ErrorPage err -> numberOfErrors++;
}
}In class Pretty:
private static String createPrettyString(Page page) {
return switch (page) {
case GitHubIssuePage issue
-> "🐈 ISSUE #" + issue.issueNumber();
case GitHubPrPage pr
-> "🐙 PR #" + pr.prNumber();
case ExternalPage ext
-> "💤 EXTERNAL: " + ext.url().getHost();
case ErrorPage err
-> "💥 ERROR: " + err.url().getHost();
};
}⇝ Simpler access with record/deconstruction patterns.
Use deconstruction patterns:
public static String createPrettyString(Page page) {
return switch (page) {
case GitHubIssuePage(
var url, var content,
int issueNumber, var links)
-> "🐈 ISSUE #" + issueNumber;
case ErrorPage(var url, var ex)
-> "💥 ERROR: " + url.getHost();
// ...
};
}⇝ Even simpler access with unnamed patterns.
Use record and unnamed patterns for simple access:
private static String createPrettyString(Page page) {
return switch (page) {
case GitHubIssuePage(_, _, int issueNumber, _)
-> "🐈 ISSUE #" + issueNumber;
case GitHubPrPage(_, _, int prNumber, _)
-> "🐙 PR #" + prNumber;
case ExternalPage(var url, _)
-> "💤 EXTERNAL: " + url.getHost();
case ErrorPage(var url, _)
-> "💥 ERROR: " + url.getHost();
};
}Looks good?
"Isn’t switching over types icky?"
Yes, but why?
⇝ It fails unpredicatbly when new types are added.
This approach behaves much better:
let’s add GitHubCommitPage implements GitHubPage
follow the compile errors!
Starting point:
record GitHubCommitPage(/*…*/) implements GitHubPage {
// ...
}Compile error because supertype is sealed.
⇝ Go to the sealed supertype.
Next stop: the sealed supertype
⇝ Permit the new subtype!
public sealed interface GitHubPage
extends SuccessfulPage
permits GitHubIssuePage, GitHubPrPage,
GitHubCommitPage {
// [...]
}Next stop: all switches that are no longer exhaustive.
private static String createPrettyString(Page page) {
return switch (page) {
case GitHubIssuePage issue -> // ...
case GitHubPrPage pr -> // ...
case ExternalPage external -> // ...
case ErrorPage error -> // ...
// missing case: GitHubCommitPage
};
}⇝ Handle the new subtype!
private static String createPrettyString(Page page) {
return switch (page) {
case GitHubIssuePage issue -> // ...
case GitHubPrPage pr -> // ...
case GitHubCommitPage page -> // ...
case ExternalPage external -> // ...
case ErrorPage error -> // ...
};
}To keep operations maintainable:
switch over sealed types
enumerate all possible types
(even if you need to ignore some)
avoid default branch
⇝ Compile error when new type is added.
operations separate from data
adding new operations is easy
adding new data types is more work,
but supported by the compiler
⇝ Like the visitor pattern, but less painful.
immutable data structures
methods (functions?) that operate on them
Isn’t this just functional programming?!
Kind of.
Functional programming:
Everything is a function
⇝ Focus on creating and composing functions.
Data-oriented programming:
Model data as data.
⇝ Focus on correctly modeling the data.
OOP is not dead (again):
valuable for complex entities or rich libraries
use whenever encapsulation is needed
still a good default on high level
DOP — consider when:
mainly handling outside data
working with simple or ad-hoc data
data and behavior should be separated
Use Java’s strong typing to model data as data:
use classes to represent data, particularly:
data as data with records
alternatives with sealed classes
use methods (separately) to model behavior, particularly:
exhaustive switch without default
pattern matching to destructure polymorphic data
| Data-Oriented Programming |
| A Lengthy Example |
| Project Amber's Arc |
All the new features we just used
were developed by Project Amber.
Profile:
project / wiki / mailing list
launched March 2017
led by Brian Goetz
In instanceof and switch, patterns can:
match against reference types
deconstruct records
nest patterns
ignore parts of a pattern
In switch:
refine the selection with guarded patterns
That (plus sealed types) are
the pattern matching basics.
This will be:
built up with more features
built out to re-balance the language
The x instanceof Y operation:
meant: "is x of type Y?"
now means: "does x match the pattern Y?"
For primitives:
old semantics made no sense
⇝ no x instanceof $primitive
new semantics can make sense
Example A: int x = 0;
x can’t literally be an instance of byte
but its value can be a byte
Example B: int y = 16_777_217;
y can’t literally be an instance of float
it can be cast to a float but not losslessly
Primitive patterns for simpler conversion checks:
int x = 0;
if (x instanceof byte b)
IO.println(b + " in [-128, 127]");
int y = 16_777_217;
if (y instanceof float f)
IO.println(f + " is a float (lossless)");(Probably sixth preview in JDK 28.)
Everything coming next is speculative,
particularly the syntax!
⚠️
Constant patterns for simpler primitive checks:
record Point(int x, int y) { }
var point = // ...
switch (point) {
// primitive pattern
case Point(var x, _) when x == 0 -> // ...
// constant pattern (probably)
case Point(0, _) -> // ...
}Currently, only records can be deconstructed.
Expand deconstruction to other types:
// ↙ INTERFACE ↙ state description
interface Point(int x, int y) {
// state description implicitly requires:
// int x();
// int y();
}
// would allow
var point = // ...
switch (point) {
case Point(var x, var y) -> // ...
}Often, there’s no need for the conditional.
Deconstruction on assignment is unconditional:
// `Point` is a type with state description
Point nextPoint() { /* ... */ }
// if you only need `x`
Point(int x, _) = nextPoint();If the type has a symmetrical construction protocol
(like records):
record Point(int x, int y) { }
var p0 = new Point(0, 0);
var p1 = p0 with { x = 1; };Patterns and deconstruction:
Other endeavors: