Design science evaluation in The Human-Machine Duet
The framework held, bent in one place, and gained five new parts
This site follows the ABCDE framework from problem to revision: how version 1 was built from the literature, how twenty practitioners in Hong Kong's music economy tested it between 21 and 25 September 2026, and how their evidence produced version 2.
Twenty participants, four coalitions
Every code on this site is clickable. Press or any other code to open its full name, wording, evidence and quotations. Full names in the top bar spells every code out in place. Keys: ← → move between sections, / finds a code.
How can creators, platforms, and policymakers in the Hong Kong music creative economy rebuild trust, attribution, and value in the age of generative artificial intelligence through an integrated socio-technical framework of provenance, remuneration, and governance?
Why design science research
The question asks for something to be built, not only explained. That places the study in the design science tradition: build an artefact that addresses a real problem, evaluate it rigorously, and report it so researchers and practitioners can both judge it (Hevner et al., 2004). Evaluation runs through the design rather than waiting for the end, and it tests usefulness where the artefact would actually be used (Sonnenberg & vom Brocke, 2012). Here the artefact is the ABCDE framework (Appendix A), and the place of use is the Hong Kong music economy after the Government's 2024 consultation on copyright and AI.
The eight steps of the study (select a step to jump to it)
1The problem
Once generative AI is in the chain, the music economy cannot reliably answer three questions.
Training data, imitation and derivation leave no reliable record of what was used, when, and on whose authority.
Value leaks through uses that existing licence categories were never designed to see.
Rules are split across copyright law, contracts, platform terms and new regulation.
The three are wired together. Weak provenance undermines payment claims, opaque payment weakens the case for reform, and fragmented governance lets both problems persist. Two of the six propositions state the problem formally: says the change is one of kind, not degree, and says no existing remedy on its own can clear all three bottlenecks.
The stakes now have a size
CISAC & PMP Strategy (2024)
Deezer (2026)
The Hong Kong opening
The Government proposed a text-and-data-mining (TDM) exception for AI training with five safeguards: lawful access, exclusion of infringing copies, record-keeping, non-availability where a licensing scheme exists, and a machine-readable opt-out. The February 2025 outcomes paper left implementation to voluntary codes of practice (CEDB & IPD, 2024, 2025). Every safeguard needs working infrastructure that the law does not specify. That gap is where the framework is designed to fit; the mapping is below.
Sixty submissions, and the missing voice
The consultation record holds 60 public written submissions. They map onto four stakeholder coalitions: rights-holders and collective-management organisations (CMOs), the technology industry, IP and legal practitioners, and creators. But the creator voice arrived only through institutions. Not one submission could be identified as the work of an individual Hong Kong composer, songwriter or performer acting outside an institutional channel.
2024 consultation: 60 written submissions
This study: 20 interviews
The literature review behind version 1
The review is integrative (Cronin & George, 2020; Snyder, 2019; Torraco, 2005; Whittemore & Knafl, 2005) and organised by concept rather than by author (Webster & Watson, 2002). Its sources are too mixed for a standard systematic review: journal articles sit beside government papers, a court judgment, technical standards and books. So source selection follows a four-stage flow adapted from PRISMA 2020 (Page et al., 2021), with tiers and domain tags in place of PRISMA's single-database assumption.
How a source got into the evidence base (Appendix B, B.3)
1Identification
Five channels: Semantic Scholar and Scite for peer-reviewed work, cross-checked in publisher catalogues; the CEDB, IPD and Legislative Council for policy; the UK Judiciary for case law; C2PA, the Content Authenticity Initiative and the Bitcoin and Ethereum white papers for standards; press catalogues for books.
2Screening
Four inclusion criteria: date (1994 to 2025), language (English or an authoritative English version), authority, and domain (music-focal, a cross-domain analogy, or a domain-neutral substrate).
3Verification gate
Three checks per record: the identifier resolves (DOI, ISBN or official URL); the authority is confirmed at the index or the issuing body; the claim cited is what the source actually says.
4Inclusion
56 records entered the source matrix (Appendix B, Table B.1), each tagged with an evidence tier and a domain.
What the 56 records are
From six arguments to six propositions and six gaps
The review is organised as six interlocking arguments. Each distils into one consolidated proposition, and the propositions became part of what the interviews tested.
| Argument in Chapter 2 | Sections | Proposition it produced |
|---|---|---|
| 1 Categorical disruption | 2.3 to 2.6 | Generative AI changes the creative economy in kind |
| 2 Insufficiency of existing remedies | 2.7 to 2.13 | No single remedy is enough |
| 3 Provenance | 2.14 to 2.19 | Provenance needs a technical record and CMO validation |
| 4 Value exchange | 2.20 to 2.24 | Smart contracts anchored in law |
| 5 Governance | 2.25 to 2.31 | Governance with people at the centre |
| 6 ABCDE integration | 2.32 to 2.38 | The cures only work together |
| Gap in the literature (Chapter 2, 2.40) | How the book answers it |
|---|---|
| Theoretical integration. Provenance, smart contracts, governance and copyright doctrine are studied as separate streams. | One architecture with explicit dependencies between the streams: . |
| Empirical Hong Kong. The literature is anchored in the US and EU; the Hong Kong consultation has not been put into practice-facing work. | Validation anchored in the Hong Kong music economy: this evaluation. |
| Cross-jurisdictional operation. Comparative law maps what each jurisdiction did, not how systems interoperate. | Interoperability as a design requirement: , , . |
| Technology-standards convergence. C2PA is mature for images and video, immature for music's layered rights. | A music-specific provenance schema and CMO-integrated metadata: . |
| Music-specific personality rights. Voice, moral rights and neighbouring rights are under-addressed. | Personality rights treated as a first-class part of provenance: . |
| Cross-disciplinary citation. Information systems, law and music scholarship rarely cite each other. | All three disciplines used as load-bearing evidence, with a verified citation tracker. |
2What success would look like
Industry-wide figures for music-rights administration are unavailable or not comparable between organisations, so success was defined through four qualitative indicators, each with a practical proxy participants could speak to.
| Indicator | What it measures | Proxy used in the interviews | Rated in this study |
|---|---|---|---|
| Whether authorship and contribution can be read end to end | Rated completeness of records; share of works with full contributor metadata | n = 11 | |
| Whether payouts can be reconciled to documented use | Time and steps to reconcile a payout; confidence that payouts reflect real shares | n = 11 | |
| Manual handoffs, disputes and time to resolve claims | Time to resolve a claim; handoffs per claim; rated friction | n = 18 | |
| Clear accountability, enforceable rules, acceptance | Rated clarity of rules; willingness to adopt; acceptance of roles | n = 18; asked of all |
The framework as a whole was also judged against the four design science criteria of Sonnenberg and vom Brocke (2012): (is it aimed at the right problem?), (would it help if built?), (are the principles consistent and specific enough?) and (could it run in Hong Kong as things stand?). Both sets of criteria were fixed before the first interview, so the evaluation could not be shaped afterwards to suit the results.
3The framework, version 1
Two layers. The theory layer claims three cures that only work together. The operational layer names the five elements any real system must use to deliver them.
The two layers (select a cure or an element for detail)
One pipeline, from creator to audit
1Capture
A digital birth certificate is created with the work: provenance at creation.
2Register
Rights and metadata are registered.
3Encode
Licensing rules, royalty splits and opt-out flags become smart contracts.
4Monitor
Usage is monitored and royalties are triggered.
5Govern
Oversight of the whole, and dispute resolution.
Twelve design principles
Written to be contestable, so participants could agree, attach conditions or reject them. Three belong to provenance, three to remuneration and six to governance.
Provenance
Recorded at creation, not rebuilt later
Verifiable without trusted intermediaries
Selective disclosure
Remuneration
Auditable payment logic
Accommodates legal change
Preserves CMO roles
Governance
Structural, relational, procedural
Co-designed with creators
Jurisdictionally adaptive
Open standards
Trust is mediated, not assumed
Voluntary, with operational gravity
Eight questions left open on purpose
Version 1 was deliberately incomplete in eight places: design choices the literature could not settle. The interviews were built to answer them.
Where the framework meets Hong Kong's five TDM safeguards
| Safeguard (CEDB & IPD, 2025) | What it demands in practice | Elements | What the interviews added |
|---|---|---|---|
| 1 Lawful access | Verifiable provenance of how a work was obtained | Supported under and | |
| 2 Infringing copies excluded | A way to identify and exclude infringing copies from training | Supported under and | |
| 3 Record-keeping and disclosure of sources | A standard, auditable, queryable record of training inputs | Supported under and | |
| 4 Not available where a licensing scheme exists | An authoritative registry of licensing schemes, and trigger logic | closed: the registry evidences availability | |
| 5 Machine-readable opt-out | A standard rights-reservation signal that the ecosystem adopts | led to covering opt-out compliance |
Working tools for participants to react to
- A music-specific metadata schema for the digital birth certificate (Appendix A, A.7)
- Ten smart-contract template clauses (A.8)
- A governance checklist for pre-deployment, operation and disputes (A.9)
- A role and responsibility matrix across the pipeline (A.10)
4Planning the evaluation
The framework makes claims about law, technology, collective management and creative practice. Only people who hold real positions in those fields can judge them, so the study used elite semi-structured interviews.
The yardstick: who counts as an elite participant
Holds an operational or institutional position in the music economy, inside one of the four coalitions.
Can react with authority to the framework's claims in their own field.
Senior staff, partners, senior associates or committee members; for creators, working composers, songwriters, producers and performers, signed or independent.
Based in Hong Kong, or running services that reach the Hong Kong market.
Individual creators over-sampled, because the consultation record had none.
What each coalition was best placed to judge (Appendix D, Table D.1)
Recruitment
Direct outreach
Through the researcher's professional networks in music, collective management, law and technology.
Referral
Each participant was asked to name two or three others holding relevant positions.
Open invitation
Organisations that made submissions to the 2024 consultation, contacted through the public record.
Planned range against interviews held
When to stop. Saturation was judged within each coalition: the point where two consecutive interviews add no new framework-relevant theme. Reviews were set at the twelfth, sixteenth and twentieth interview, with a minimum of 16 and a cap of 22 (Appendix D, D.14.3).
The interview guide
A common core for everyone, then one module per coalition
Core questions, all 20 participants
Q1 role and perspective. Q2 the most important problem. Q3 the three cures. Q4 the ABCDE elements. Q5 the twelve design principles. Q6 the Government's five TDM safeguards. Q7 the condition for a voluntary code to succeed. Q8 anything missed.
Three revision rules, written before fieldwork
Appendix D fixed how findings would turn into revisions before anyone was interviewed. Because the rules came first, the study could not quietly promote a convenient finding or bury an awkward one. "Majority" means more than half of a coalition: 4 of 6 creators, 3 of 5 CMO, 3 of 4 technology, or 3 of 5 legal participants.
Ethics
Approved by Golden Gate University's Institutional Review Board; the approved papers are fixed. Participants signed an informed-consent form and are identified only by role-based codes. They may withdraw at any time without penalty and, after the interview, within seven days by email, with their data destroyed immediately. They may review the record of their interview within fourteen days of receiving it, and any direct quotation goes back to them for approval before publication. Records are kept securely for three years and then destroyed; raw data are seen only by the researcher and the dissertation committee.
5Fieldwork
Twenty interviews, all in person, over five days in September 2026. Nineteen in English, one in Cantonese.
Interviews per day, by coalition
Who took part (whole sample)
What the record is
- The record of each interview is the set of notes the researcher captured live in the Interview Desk, a web application built for the study, together with the ratings. It is not a verbatim transcript.
- In one case (C4-07, Question 6) a line was added from the researcher's recollection after the interview. It is dated 25 September 2026 and flagged, so it can be told apart from what was recorded live.
6Analysis
Thematic synthesis (Thomas & Harden, 2008): coding against the framework, then open coding for what the framework missed, with every decision checked by the researcher and verified in MAXQDA.
From interview notes to verified counts
1Interview notes
20 records captured live in the Interview Desk.
2Coded documents
Four coalition files in MAXQDA 26; every line tagged with its question; ratings held as variables.
3Run 1: deductive
Each line coded to the framework element it speaks to, plus a stance.
4Researcher review
All 629 lines read; 134 decisions with reasons; 10 lines removed.
5Run 2: inductive
Open reading for what the framework misses; 8 candidate themes reviewed into 7 themes and 1 sub-theme.
6Verification
MAXQDA Code Matrix Browser, one count per person; matched the coded data exactly.
Run 1: does the evidence fit the framework?
The codebook is the framework itself: the four indicators, the three cures and twelve principles, the five elements, the eight open questions, the six propositions, the four evaluation criteria, and markers for quotation candidates. Each line got a stance: Support, Qualification or Criticism. Six rules were settled during the review and applied to every line.
- A stance is recorded only on framework elements. Open questions and new themes carry none: there is nothing yet to agree or disagree with.
- A problem the framework is designed to fix counts as Support for the element that fixes it.
- Agreement plus a suggestion, or "this would be better with law behind it", counts as Support.
- Agreement that depends on a condition ("only if", "unless") counts as Qualification.
- Answers about the Government's five safeguards (Q6) are coded to and , not to : they concern the Government's proposal, not the framework's voluntary design.
- A creator saying the disruption has not arrived yet, or has arrived as a gain, counts as Qualification of .
What the line-by-line review changed (134 decisions)
Final stance codings
Run 2: what does the evidence say that the framework does not?
A second, open reading looked for material no framework code could hold. It found eight candidate themes. Review corrected seven misplaced lines, folded a patent-style searchable registry into as the sub-theme (two of its three people were really making a point about government), renamed , and marked as minor. The result is seven themes and one sub-theme; there is no IND-004 in the final list.
Verification
Apart from the line totals above, every count in the findings is a count of people, not lines: someone who made a point three times counts once. The coded project was exported from MAXQDA's Code Matrix Browser with coalitions as columns and each document counted once, then compared with counts computed from the coded data. Every theme row and the Support and Qualification split matched.
7What held
All twenty participants supported at least one framework element, twelve attached a condition somewhere, and nobody rejected anything.
Support for each design principle and proposition, by coalition
The strongest results sit where everyone was asked. and were supported by all twenty participants, and by eighteen. Among the principles, was supported by eighteen, by sixteen and by fifteen. had every creator and every CMO participant behind it, which matters because the framework asks both groups to trust the same infrastructure.
Trigger 1 not fired No principle was rejected by anyone, let alone by a majority of two coalitions ().
Where it bent: voluntary adoption comes with conditions
carried the most conditions: five people, at least one in every coalition. None rejected voluntary adoption. Each named what it would take.
Read together, they describe an order of events. A second legal participant put it directly: best practice first, and amend the law at the end. So the finding confirms and adds one sentence to it.
Smaller, more technical conditions
- : four people named what smart contracts need to work at scale: legal support, a clear basis for sharing costs, government advocacy, enough market scale, or faster blockchain throughput.
- : two technology participants warned that co-design could be captured by larger players, or that a fair mix of participants is hard to define.
- : two feared tighter rules would hurt Hong Kong's competitiveness or push AI training elsewhere.
- Opt-out under : three would use an opt-out only if it were monitored, carried no extra charge, or allowed exemptions such as education.
- : three creators said the disruption had not reached them yet, or had reached them as a gain in their own production.
What the ratings add
Where participants gave figures, four ratings were common enough to show. Each dot is one person.
The current system creates a lot of friction (, median 4.5) and little confidence that payments are right (, median 2). The framework's rules read as clear to the people who would have to follow them (, median 4). Record completeness () was spread from 1 to 5.
Would they join a voluntary code built on the framework? ()
Which open questions the evidence settled
Trigger 2 fired An open question becomes a design choice when a majority of one coalition answers it the same way (). Bars show how many people in each coalition addressed the question.
What the framework missed
Trigger 3 fired A new theme raised by a majority of one coalition is considered for the design (). The tick on each bar marks half the coalition; a majority must pass it.
Five themes cleared the bar and changed the design. is the largest: thirteen people, including all five legal and four of five CMO participants. Eight raised it unprompted when asked for the most important problem, and those eight came from all four coalitions. came from twelve people with no agreement at all on who should pay. was a majority view among creators, among legal participants, and among CMO participants, who doubted anyone could check whether an opt-out had been honoured.
Two did not change the design. clears the bar technically but explains why the framework matters, not how it should work, so it enters the rationale for . , from three people, is recorded as a refinement of the smart-contract layer.
8Version 2
Fourteen changes, adopted on 26 September 2026, each traced to the rule that triggered it and the evidence behind it. No principle was removed.
Confirmed
- one sentence added: voluntary first, law later as a complementary layer
- All other principles held; none rejected or removed
Open questions closed
- disclosure to legitimate-interest parties through permissioned access
- the registry evidences licence availability
- voice and style registered as a provenance attribute
- technical part only
Added or extended
- new: human–AI authorship threshold
- new: who pays for provenance
- education and ethical-readiness strand
- extended to opt-out compliance
- government as convener; IPD as possible registry host
Still open
- tribunal or voluntary resolution (narrowed)
- adoption incentives (narrowed)
- rights split across jurisdictions
- unidentified rights-holders (not addressed)
| # | Item | Change | Basis |
|---|---|---|---|
| 1 | Confirmed, with one sentence added: voluntary first, statutory backing later as a complementary layer | Most-qualified principle (5 people, 4 coalitions); "best practice first" | |
| 2 | Closed: voice and style registered as a provenance attribute; legal treatment left to future reform | : CMO 4 of 5 | |
| 3 | Closed: the registry evidences licence availability; the trigger works only with a legal basis or auditable monitoring | : CMO 4 of 5 | |
| 4 | Closed: disclosure to legitimate-interest parties through permissioned access; access standard set jointly | : Tech 4 of 4 | |
| 5 | Split: technical part closed; offshore enforcement stays open | technical part: Tech 4 of 4 | |
| 6 | Open, narrowed: anchor-adopter approach recorded as the leading direction | Raised widely, without convergence | |
| 7 | Added as a new open question | : | |
| 8 | Added as a new open question | : | |
| 9 | Extended with an education and ethical-readiness strand | : | |
| 10 | Extended to opt-out compliance: an auditable record that opt-out signals were honoured | : | |
| 11 | Note added: government as convener, not first-mover legislator; IPD as a possible registry host | : , | |
| 12 | Rationale only: the distinct value of human creativity | ||
| 13 | No change; creator-configurable terms recorded as a refinement | (3 people) | |
| 14 | , , | Remain open; OQ3 narrowed; OQ7 reported as a gap in the interview guide | No convergence, or no data |
Two of these changes were anticipated when version 1 was written: Appendix A named a refinement of and an extension toward voice and likeness as the most likely revisions. The evaluation produced both, with the direction coming from participants rather than the draft. The five additions (items 7 to 11) were not anticipated at all. They are the clearest sign that the interviews did more than confirm what was already written.
Voices from the interviews
Quotations awaiting participant approval. Wording is from the researcher's live notes, edited for grammar only. As the approved protocol requires, each quotation is being returned to its participant before publication. Participants are identified by coalition and code only.
How far the results can be trusted
Rigour in qualitative work comes from checks built into the process, not from judging the work afterwards (Morse et al., 2002). Five were built in.
What makes them credible
- Indicators, criteria and revision rules were fixed before the first interview.
- All 629 lines were read; each of the 134 decisions carries a written reason that can be audited.
- Contrary evidence was kept and re-read. Dropping one legal participant whose views diverged was considered and rejected, because removing a dissenting voice would bias the result.
- All counts are counts of people, verified against an independent MAXQDA export.
- Every change to version 2 is traced to its rule and its evidence.
What limits them
- The record is the researcher's live notes, not a verbatim transcript.
- The fourteen-day participant review is under way; corrections will be logged.
- Support is very high (403 against 28). Some may be politeness, a risk raised by recruiting through the researcher's own networks. About half the Supports added in review are problem statements rather than endorsements, which reduces the risk without removing it.
- One jurisdiction only, so claims about rest on participants' views, not tests abroad.
- Coalitions answered different modules; cross-coalition comparisons hold only for shared questions.
- Ratings come from 11 to 18 people and are descriptive.
- One legal participant was added after saturation, and was not tested.
What happens next
Design science treats an artefact as never finished: version 2 is the starting point of the next cycle (Hevner et al., 2004; Sonnenberg & vom Brocke, 2012). Four steps are set: participants complete their fourteen-day review; each quotation goes back to its source for approval; the committee reviews version 2; and the book carries version 2. The agenda for the next cycle is , , and , the questions most likely to decide whether a voluntary framework can work in practice.
Code key
Every label used on this site. Select a row for the full entry.
Sources
Works cited on this site, as verified in the book's reference list (APA 7th edition). Book locations such as "Appendix A, A.12" refer to The Human-Machine Duet.
- Commerce and Economic Development Bureau & Intellectual Property Department. (2024, July). Public consultation on copyright and artificial intelligence. Government of the Hong Kong Special Administrative Region. https://www.ipd.gov.hk/filemanager/ipd/en/share/consultation-papers/Eng-Copyright-and-AI-Consultation-Paper-20240708.pdf
- Commerce and Economic Development Bureau & Intellectual Property Department. (2025, February 18). Enhancement of the Copyright Ordinance regarding protection for artificial intelligence technology development: Outcomes of public consultation and proposed way forward (LC Paper No. CB(2)240/2025(04)). Legislative Council Panel on Commerce, Industry, Innovation and Technology. https://www.legco.gov.hk/yr2025/english/panels/ci/papers/ci20250218cb2-240-4-e.pdf
- CISAC & PMP Strategy. (2024, December). Study on the economic impact of generative AI in the music and audiovisual sectors. CISAC. https://www.cisac.org/services/reports-and-research/cisacpmp-strategy-ai-study
- Cronin, M. A., & George, E. (2020). The why and how of the integrative review. Organizational Research Methods, 26(1), 168–192. https://doi.org/10.1177/1094428120935507
- Deezer. (2026, April 20). AI-generated tracks now represent 44% of all new uploaded music [Press release]. https://newsroom-deezer.com/2026/04/ai-generated-tracks-represent-44-of-new-uploaded-music/
- Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625
- Morse, J. M., Barrett, M., Mayan, M., Olson, K., & Spiers, J. (2002). Verification strategies for establishing reliability and validity in qualitative research. International Journal of Qualitative Methods, 1(2), 13–22. https://doi.org/10.1177/160940690200100202
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
- Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, 333–339. https://doi.org/10.1016/j.jbusres.2019.07.039
- Sonnenberg, C., & vom Brocke, J. (2012). Evaluations in the science of the artificial: Reconsidering the build-evaluate pattern in design science research. In K. Peffers, M. Rothenberger, & B. Kuechler (Eds.), Design Science Research in Information Systems. Advances in Theory and Practice (DESRIST 2012, LNCS 7286, pp. 381–397). Springer. https://doi.org/10.1007/978-3-642-29863-9_28
- Thomas, J., & Harden, A. (2008). Methods for the thematic synthesis of qualitative research in systematic reviews. BMC Medical Research Methodology, 8, Article 45. https://doi.org/10.1186/1471-2288-8-45
- Torraco, R. J. (2005). Writing integrative literature reviews: Guidelines and examples. Human Resource Development Review, 4(3), 356–367. https://doi.org/10.1177/1534484305278283
- Webster, J., & Watson, R. T. (2002). Analyzing the past to prepare for the future: Writing a literature review. MIS Quarterly, 26(2), xiii–xxiii. https://www.jstor.org/stable/4132319
- Whittemore, R., & Knafl, K. (2005). Integrative review: Updated methodology. Journal of Advanced Nursing, 52(5), 546–553. https://doi.org/10.1111/j.1365-2648.2005.03621.x