machi

greg/machi

Author	SHA1	Message	Date
Scott Lystig Fritchie	fe8ff6033d	Make better state transition choices in AP mode	2015-09-11 19:14:41 +09:00
Scott Lystig Fritchie	a0c129c16d	Bugfix: wow, a chain state transition sanity check bug	2015-09-11 17:32:52 +09:00
Scott Lystig Fritchie	8df7d58365	Add partition simulator support to fitness service	2015-09-11 16:45:29 +09:00
Scott Lystig Fritchie	efe6ce7894	WIP: small refactoring to prepare for fitness server 'use' of partition simulator	2015-09-11 16:03:49 +09:00
Scott Lystig Fritchie	35e8efeb96	Add timer:sleep() to accomodate machi_chain_manager1_converge_demo	2015-09-11 15:56:02 +09:00
Scott Lystig Fritchie	bbf925d132	Add fault injection method via C100 to test C103 admin down cycle	2015-09-10 18:05:55 +09:00
Scott Lystig Fritchie	41737ae62a	Add delete_admin_down API implementation, oops!	2015-09-10 18:05:18 +09:00
Scott Lystig Fritchie	d45c249e89	Add admin down status API to fitness server	2015-09-10 17:30:11 +09:00
Scott Lystig Fritchie	c14b9ce50f	Minor cleanup, add more partitions to converge demo	2015-09-10 16:39:15 +09:00
Scott Lystig Fritchie	af94d1c1c3	Bugfix: ExpectedUPI error in A40	2015-09-10 02:15:49 +09:00
Scott Lystig Fritchie	daf3a3d65a	Remove some verbose debugging cruft	2015-09-10 01:47:46 +09:00
Scott Lystig Fritchie	329a5e0682	Bugfix: damn, no idea how many problems this 5 month old bug caused	2015-09-10 01:33:55 +09:00
Scott Lystig Fritchie	5943494d54	Add ExpectedUPI to A40's AmHosedP clause	2015-09-10 00:43:37 +09:00
Scott Lystig Fritchie	10c655ebfe	WIP: fix one source of problems, now shift back to 'TODO this clause needs more review'	2015-09-09 23:59:40 +09:00
Scott Lystig Fritchie	b7aa33c617	Yeah, nearly there. AP fails occasionally in multiple-asymmetric-partition sequence	2015-09-09 23:10:39 +09:00
Scott Lystig Fritchie	72141c8ecb	WIP: split A30 into A30/A31 based on AllHosed	2015-09-09 21:06:40 +09:00
Scott Lystig Fritchie	5029911b52	WIP: remove verbose goop	2015-09-09 20:46:52 +09:00
Scott Lystig Fritchie	38ea36fc1c	WIP: Stand back, I'm going to try math! ... It works, {redacted}!	2015-09-09 20:45:57 +09:00
Scott Lystig Fritchie	27891bc5e9	WIP: 'broadcast'/spam works! async reminder ticks remain!	2015-09-09 19:14:52 +09:00
Scott Lystig Fritchie	dd095f117f	Derp, fix smoke_test() for machi_fitness:map_set()	2015-09-09 16:49:27 +09:00
Scott Lystig Fritchie	21015efcbb	WIP: Stand back, I'm going to try CRDTs!	2015-09-08 19:13:03 +09:00
Scott Lystig Fritchie	7af863d840	Add stubs of machi_fitness server	2015-09-08 16:13:07 +09:00
Scott Lystig Fritchie	185c9eb313	WIP: add failing eunit placeholder for spam	2015-09-07 15:38:23 +09:00
Scott Lystig Fritchie	c7684f660c	WIP: Friday evening/Monday morning, laying groundwork for spam "broadcast"	2015-09-07 15:20:10 +09:00
Scott Lystig Fritchie	4376ce9ec1	Remove all flap counting and inner projection stuff	2015-09-04 17:17:49 +09:00
Scott Lystig Fritchie	42aeecd9db	Fix machi_projection_store_test error	2015-09-04 15:24:16 +09:00
Scott Lystig Fritchie	3c1026da28	WIP: too tired to continue tonight	2015-09-01 22:10:45 +09:00
Scott Lystig Fritchie	4378ef7b54	Bugfix: inner->outer proj @ A30	2015-09-01 00:51:46 +09:00
Scott Lystig Fritchie	e79265228e	Bugfix: more correct for inner->outer sanity transition	2015-08-31 22:14:28 +09:00
Scott Lystig Fritchie	1e5d58b22d	Bugfix: more to ignore in make_basic_comparison_stable()	2015-08-31 17:57:37 +09:00
Scott Lystig Fritchie	bce225a200	Bugfix: a30_make_inner_projection() ignore newprop down list if none proj	2015-08-31 17:03:12 +09:00
Scott Lystig Fritchie	a095e0cfc3	Bugfix: ignore creation_time in make_comparison_stable()	2015-08-31 15:40:19 +09:00
Scott Lystig Fritchie	c637939cc2	Bugfix: A29 should trigger if EpochID (not Epoch# alone) differs	2015-08-31 15:21:17 +09:00
Scott Lystig Fritchie	5422dc45c2	Bugfix: derp in A29 revival	2015-08-31 14:44:05 +09:00
Scott Lystig Fritchie	004c686c8c	WIP: remove make_zerf() from calc_projection(); add make_zerf() to resurrected A29. Status: broken, needs work	2015-08-30 20:39:58 +09:00
Scott Lystig Fritchie	a449025e8b	Bugfix: epoch handling around none proj: epoch 0 only at first bootstrap!	2015-08-30 19:53:47 +09:00
Scott Lystig Fritchie	ec2e7b5669	Sunday experiment: all-but-remove A29, feels right but definitely not sure yet	2015-08-30 16:08:14 +09:00
Scott Lystig Fritchie	0dc53274d1	Get more aggressive about AllHosed+down nodes for inner proj	2015-08-30 02:22:59 +09:00
Scott Lystig Fritchie	771164b82f	Bugfix: Flapping manifesto, leaving #2 : only if not me	2015-08-30 00:50:23 +09:00
Scott Lystig Fritchie	4b83893047	Bugfix: minor flap count bookeeping error	2015-08-30 00:50:03 +09:00
Scott Lystig Fritchie	a7db3a26c6	Bugfix: a30_make_inner_projection() compatible inner if not none proj	2015-08-30 00:04:13 +09:00
Scott Lystig Fritchie	53d865b247	Bugfix: serious derp fix for A30's inner->outer	2015-08-29 23:42:47 +09:00
Scott Lystig Fritchie	5c8b255da9	Bugfix: first new CP experiments with chain len=5	2015-08-29 22:40:18 +09:00
Scott Lystig Fritchie	94394d3429	Bugfix: allow none proj to re-emerge from flapping (more) See comments added in this commit at A40. So far, I've been doing CP mode testing with a handful of (very useful) network partition combinations using: machi_chain_manager1_converge_demo:t(3, [{private_write_verbose,true}, {consistency_mode, cp_mode}, {witnesses, [a]}]). Next steps: * Expand number & types of partitions * Expand to chain lengths of 5 and beyond	2015-08-29 21:36:53 +09:00
Scott Lystig Fritchie	ee19a0856b	WIP: justincase	2015-08-29 19:59:46 +09:00
Scott Lystig Fritchie	6b84cd6e6a	Reduce poll sleep time when running with partition simulator	2015-08-29 18:30:53 +09:00
Scott Lystig Fritchie	dc5ae4047a	Bugfix: react_to_env_A30 inner->norm fix, make_zerf() none proj derp fix	2015-08-29 18:01:13 +09:00
Scott Lystig Fritchie	c9340a662d	Bugfix: force stable creation_time on inner none proj	2015-08-29 15:06:57 +09:00
Scott Lystig Fritchie	6d9526b379	Add more ?REACT()	2015-08-29 13:13:31 +09:00
Scott Lystig Fritchie	f21fcdd7be	Bugfix: none proj must flap, undo previous commits, which may cause mess later	2015-08-29 13:13:23 +09:00
Scott Lystig Fritchie	af0ade9840	Bugfix: projection checksum fix in A30	2015-08-29 12:33:41 +09:00
Scott Lystig Fritchie	582f9e5eab	Bugfix: fix effectively-none-projection transition to C100. Still buggy	2015-08-28 23:08:38 +09:00
Scott Lystig Fritchie	403cb5b7a6	WIP: improvements, but now flapping inner epoch keeps increasing {sigh}	2015-08-28 21:13:54 +09:00
Scott Lystig Fritchie	9edd91f48e	Bugfixes for a->b column transition & flap dampening	2015-08-28 20:06:09 +09:00
Scott Lystig Fritchie	18aac6e489	WIP: undo AmFlappingNow_p condition added at commit `3dfe5c2`	2015-08-28 18:39:18 +09:00
Scott Lystig Fritchie	3dfe5c2677	WIP: fix annotation history on disk	2015-08-28 18:37:11 +09:00
Scott Lystig Fritchie	8ca1ffdb13	WIP: bugfixes and lots of verbose goop added	2015-08-28 01:55:31 +09:00
Scott Lystig Fritchie	deb2cdee2c	Bugfix: correct epoch number checking when inner proj	2015-08-27 22:22:15 +09:00
Scott Lystig Fritchie	93b9b948fc	WIP: debugging, uff da	2015-08-27 22:02:23 +09:00
Scott Lystig Fritchie	efb89efb0d	Reduce verbosity	2015-08-27 20:27:33 +09:00
Scott Lystig Fritchie	0eaa008810	Change checksum algorithm to exclude 'flap' also	2015-08-27 20:27:24 +09:00
Scott Lystig Fritchie	12b74a52fd	WIP: pre-dinner paranoid checkin	2015-08-27 18:45:27 +09:00
Scott Lystig Fritchie	65cd18939c	WIP: changes to annotation management	2015-08-27 17:58:43 +09:00
Scott Lystig Fritchie	8a61a85ae0	WIP: rewrite make_zerf() to use new annotation scheme	2015-08-27 16:19:22 +09:00
Scott Lystig Fritchie	28335a1310	Add CP mode unwedge. All eunit tests are passing again.	2015-08-26 18:47:39 +09:00
Scott Lystig Fritchie	9222881689	Oops, bugfixes	2015-08-26 17:51:43 +09:00
Scott Lystig Fritchie	568e165f4f	Allow pstore -> FLU unwedge only in ap_mode, machi_cr_client_test broken (uses cp_mode)	2015-08-26 15:51:14 +09:00
Scott Lystig Fritchie	e8f3ab381d	Add set_consistency_mode() to projection store API, use it	2015-08-26 14:57:51 +09:00
Scott Lystig Fritchie	c0ee323637	Our new unit test works, yay	2015-08-25 19:42:33 +09:00
Scott Lystig Fritchie	83f49472db	WIP: intermediate refactoring	2015-08-25 19:31:05 +09:00
Scott Lystig Fritchie	6dbe887298	Remove old cruft, including hugly HTTP server hack	2015-08-25 18:49:48 +09:00
Scott Lystig Fritchie	1c5a17b708	WIP: adjust throttle of flapping 'shut up'	2015-08-25 17:01:14 +09:00
Scott Lystig Fritchie	9a86453753	WIP: half-baked idea, stopping for the night (more) So, I'm 50% sure this is a good idea for CP mode: if there's a later public projection than P_current, then who knows what we might have missed. So, call make_zerf() to find out the absolute latest. Problem: flapping state appears to be lost, booo.	2015-08-24 21:54:30 +09:00
Scott Lystig Fritchie	ea61fe78bf	Add flap disabler for 3 seconds after up/down change	2015-08-24 20:38:54 +09:00
Scott Lystig Fritchie	2f82fe0487	WIP: cp_mode improvements	2015-08-24 19:04:26 +09:00
Scott Lystig Fritchie	66cafe066e	Remove proj_i_history, tweak AllAreFlapping_and_IamBad_and_NotRelevant_p in B10	2015-08-23 20:47:43 +09:00
Scott Lystig Fritchie	70022d11ce	Add damper check for flapping of inner projections, whee!	2015-08-23 20:00:19 +09:00
Scott Lystig Fritchie	561e60a7ac	WIP: start adding support to detect flapping of inner projections (ha!)	2015-08-23 17:50:25 +09:00
Scott Lystig Fritchie	0136fccff7	CP mode fix a30_make_inner_projection	2015-08-23 16:43:15 +09:00
Scott Lystig Fritchie	2d050ff7a6	Fix ?REACT() FSM names: a30->a40	2015-08-23 15:46:57 +09:00
Scott Lystig Fritchie	34d35fab63	Shorten the verbose output of private_write_verbose	2015-08-22 23:30:30 +09:00
Scott Lystig Fritchie	51a06844d5	Fix epoch number reuse bug when transiting C103	2015-08-22 21:40:21 +09:00
Scott Lystig Fritchie	0414da783a	Fix repairs when everyone is in stable flapping state	2015-08-22 21:27:01 +09:00
Scott Lystig Fritchie	a0477d62c0	WIP: bugfix for checking latest proj's flap count	2015-08-22 14:50:10 +09:00
Scott Lystig Fritchie	0278d7254b	Add A29 state for shouting circuit breaker for long long loops	2015-08-20 23:04:27 +09:00
Scott Lystig Fritchie	b46730eb2c	WIP: adjust the flapping manifest: delete clause 3	2015-08-20 21:28:56 +09:00
Scott Lystig Fritchie	71decc5dc0	WIP: AP mode less bad again	2015-08-20 18:47:50 +09:00
Scott Lystig Fritchie	4e7d1f2310	WIP: egadz, a refactoring mess, but finally AP mode not sucky	2015-08-20 17:32:46 +09:00
Scott Lystig Fritchie	a71e9543fe	WIP: refactoring inner handling, but ... (more) There are a couple of weird things in the snippet below (AP mode): 22:32:58.209 b uses inner: [{epoch,136},{author,c},{mode,ap_mode},{witnesses,[]},{upi,[b,c]},{repair,[]},{down,[a]},{flap,undefined},{d,[d_foo1,{ps,[{a,b}]},{nodes_up,[b,c]}]},{d2,[]}] (outer flap epoch 136: {flap_i,{{{epk,115},{1439,904777,11627}},28},[a,{a,problem_with,b},{b,problem_with,a}],[{a,{{{epk,126},{1439,904777,149865}},16}},{b,{{{epk,115},{1439,904777,11627}},28}},{c,{{{epk,121},{1439,904777,134392}},15}}]}) (my flap {{epk,115},{1439,904777,11627}} 29 [{a,{{{epk,126},{1439,904777,149865}},28}},{b,{{{epk,115},{1439,904777,11627}},29}},{c,{{{epk,121},{1439,904777,134392}},26}}]) 22:32:58.224 c uses inner: [{epoch,136},{author,c},{mode,ap_mode},{witnesses,[]},{upi,[b,c]},{repair,[]},{down,[a]},{flap,undefined},{d,[d_foo1,{ps,[{a,b}]},{nodes_up,[b,c]}]},{d2,[]}] (outer flap epoch 136: {flap_i,{{{epk,115},{1439,904777,11627}},28},[a,{a,problem_with,b},{b,problem_with,a}],[{a,{{{epk,126},{1439,904777,149865}},16}},{b,{{{epk,115},{1439,904777,11627}},28}},{c,{{{epk,121},{1439,904777,134392}},15}}]}) (my flap {{epk,121},{1439,904777,134392}} 28 [{a,{{{epk,126},{1439,904777,149865}},28}},{b,{{{epk,115},{1439,904777,11627}},28}},{c,{{{epk,121},{1439,904777,134392}},28}}]) CONFIRM by epoch inner 136 <<103,64,252,...>> at [b,c] [] Priv1 [{a,{{132,<<"Cï\|ÿzKX:Á"...>>},[a],[c],[b],[],false}}, {b,{{127,<<185,139,3,2,96,189,...>>},[b,c],[],[a],[],false}}, {c,{{133,<<145,71,223,6,177,...>>},[b,c],[a],[],[],false}}] agree false Pubs: [{a,136},{b,136},{c,136}] DoIt, 1. Both the "uses inner" messages and also the "CONFIRM by epoch inner 136" show that B & C are using the same inner projection. However, the 'Priv1' output shows b & c on different epochs, 127 & 133. Weird. 2. I've added an infinite loop, probably in this commit. :-(	2015-08-18 22:35:57 +09:00
Scott Lystig Fritchie	9bf0eedb64	WIP: add the flapping manifesto, much is muchmuch better now	2015-08-18 20:49:36 +09:00
Scott Lystig Fritchie	e9268080af	Finish/catchup commit from end of last week, silly me	2015-08-17 20:14:29 +09:00
Scott Lystig Fritchie	48e82ac1a4	WIP: use digraph to calculate better AllHosed	2015-08-14 22:29:20 +09:00
Scott Lystig Fritchie	20f2bf4b92	WIP: more ?REACT() tracing	2015-08-14 22:28:50 +09:00
Scott Lystig Fritchie	d2ce8f8447	Fix repair bug that has survived witness additions, oops	2015-08-14 19:30:36 +09:00
Scott Lystig Fritchie	9e02a1ea73	Add more ?REACT() tracing	2015-08-14 19:30:05 +09:00
Scott Lystig Fritchie	5aff775383	WIP: it's ugly, but CP+witnesses is mostly working?	2015-08-14 17:05:16 +09:00
Scott Lystig Fritchie	4e66d7bd91	WIP: keep CMode propagation consistent, but still violating CP transition safety	2015-08-14 00:12:13 +09:00
Scott Lystig Fritchie	14fad2d704	End-to-end chain state checking is still broken (more) If we use verbose output from: machi_chain_manager1_converge_demo:t(3, [{private_write_verbose,true}, {consistency_mode, cp_mode}, {witnesses, [a]}]). And use: tail -f typescript_file \| egrep --line-buffered 'SET\|attempted\|CONFIRM' ... then we can clearly see a chain safety violation when moving from epoch 81 -> 83. I need to add more smarts to the safety checking, both at the individual transition sanity check and at the converge_demo overall rolling sanity check. Key to output: CONFIRM by epoch {num} {csum} at {UPI} {Repairing} SET # of FLUs = 3 members [a,b,c]). CONFIRM by epoch 1 <<96,161,96,...>> at [a,b] [c] CONFIRM by epoch 5 <<134,243,175,...>> at [b,c] [] CONFIRM by epoch 7 <<207,93,225,...>> at [b,c] [] CONFIRM by epoch 47 <<60,142,248,...>> at [b,c] [] SET partitions = [{c,b},{c,a}] (1 of 2) at {22,3,34} CONFIRM by epoch 81 <<223,58,184,...>> at [a,b] [] SET partitions = [{b,c},{b,a}] (2 of 2) at {22,3,38} CONFIRM by epoch 83 <<33,208,224,...>> at [a,c] [] SET partitions = [] CONFIRM by epoch 85 <<173,179,149,...>> at [a,c] [b]	2015-08-13 22:16:28 +09:00
Scott Lystig Fritchie	f7121f8845	Witness + flapping seems to mostly work, yay!	2015-08-13 21:24:56 +09:00
Scott Lystig Fritchie	425b9c8f60	Merge slf/projection-conditional-write branch	2015-08-13 19:10:48 +09:00
Scott Lystig Fritchie	dcbc3b45ff	C110: handle proj store private write failure when conditional fails	2015-08-13 18:45:15 +09:00
Scott Lystig Fritchie	9768f3c035	Projection store private write returns bad_arg if max_public_epochid is greater	2015-08-13 18:44:25 +09:00
Scott Lystig Fritchie	58d840ef7e	Minor react changes, minor fix for return val of A50	2015-08-13 18:43:41 +09:00
Scott Lystig Fritchie	d4275e5460	WIP: zerf_find_last_common() fix, eunit passes & very basic len=3 converge demo works	2015-08-13 15:41:18 +09:00
Scott Lystig Fritchie	0b8de235a9	WIP: zerf_find_last_common(), but is confused/broken by partial write @ private	2015-08-13 14:21:31 +09:00
Scott Lystig Fritchie	054397d187	WIP: find last common majority epoch	2015-08-12 17:53:39 +09:00
Scott Lystig Fritchie	d340b6a706	WIP: Duh, fix think-o in a40_latest_author_down()	2015-08-12 17:37:45 +09:00
Scott Lystig Fritchie	8e2a688526	WIP: cp_mode code from last Friday	2015-08-11 15:24:26 +09:00
Scott Lystig Fritchie	512251ac55	Adjust flap_limit constant	2015-08-07 12:29:10 +09:00
Scott Lystig Fritchie	3ca0f4491d	WIP: always start chain manager with none projection	2015-08-06 19:24:14 +09:00
Scott Lystig Fritchie	0d7f6c8d7e	WIP: chain transitions are now fully (?) aware of witness servers	2015-08-06 17:48:31 +09:00
Scott Lystig Fritchie	e9c4e2f98d	WIP: rearrange CP mode projection calc	2015-08-06 15:22:04 +09:00
Scott Lystig Fritchie	82b6726261	Revert UPI [] -> [FirstRepairing] to commit `91496c6`	2015-08-06 15:21:44 +09:00
Scott Lystig Fritchie	01da7a7046	TODO WTF was I thinking here??....	2015-08-06 14:13:19 +09:00
Scott Lystig Fritchie	dcf532bafd	WIP: Witness test expansion	2015-08-05 18:23:44 +09:00
Scott Lystig Fritchie	0f18ab8d20	Add better (?) timeout handling to machi_cr_client.erl gen_server calls	2015-08-05 17:48:06 +09:00
Scott Lystig Fritchie	e3d9ba2b83	WIP: Witness test expansion	2015-08-05 17:17:25 +09:00
Scott Lystig Fritchie	b21803a6c6	Fix witness calculation projections, part II	2015-08-05 16:05:03 +09:00
Scott Lystig Fritchie	f43a5ca96d	Fix witness calculation projections, part I	2015-08-05 15:50:32 +09:00
Scott Lystig Fritchie	91496c656b	Oops, fix PB stuff to add witnesses	2015-08-05 12:53:20 +09:00
Scott Lystig Fritchie	3f51357577	WIP: pre-travel code, not sure if good, check in for history	2015-07-30 13:12:08 -07:00
Scott Lystig Fritchie	aa1a31982a	Add 'witnesses' to machi_projection:make_summary()	2015-07-30 13:11:43 -07:00
Scott Lystig Fritchie	6e521700bd	WIP: Adding witness_smoke_test_ but it's broken (more) So, the problem is that the chain manager isn't finishing repair because UPI=[a], and a is a witness, and a can't do the list files etc etc repair stuff that repairer FLUs need to do. The best (?) way forward is to add some advance smarts to the chain manager so that it doesn't propose a UPI of 100% witnesses?	2015-07-21 19:05:04 +09:00
Scott Lystig Fritchie	432190435e	Add witness_mode to FLU	2015-07-21 17:29:33 +09:00
Scott Lystig Fritchie	88d3228a4c	Fix various problems with repair not being aware of inner projections	2015-07-20 16:25:42 +09:00
Scott Lystig Fritchie	9ae4afa58e	Reduce chmgr verbosity a bit	2015-07-20 14:58:21 +09:00
Scott Lystig Fritchie	e14493373b	Bugfix: add missing reset of not_sanes dictionary, fix comments	2015-07-20 14:04:25 +09:00
Scott Lystig Fritchie	f7ef8c54f5	Reduce # of assumptions made by ch_mgr + simulator for 'repair_airquote_done'	2015-07-19 13:32:55 +09:00
Scott Lystig Fritchie	b8c642aaa7	WIP: bugfix for rare flapping infinite loop (done^2 fix I hope) How can even computer? So, there's a flavor of the flapping infinite loop problem that can happen without flapping being detected (by the existing flapping detector, that is). That detector relies on a series of accepted projections to converge to a single projection repeated X times. However, it's possible to have a race with a simulated repair "finishing" that causes a problem so that no more projections are ever accepted. Oops. See also: new comments in do_react_to_env().	2015-07-19 00:43:10 +09:00
Scott Lystig Fritchie	57b7122035	Fix bug found by PULSE that's not directly chain manager-related (more) PULSE managed to create a situation where machi_proxy_flu_client1 would appear to fail a remote attempt to write_projection. The client would retry, but the 1st attempt really did get through to the server. So, if we hit this case, we try to read the projection, and if it's exactly equal to what we tried to write, we consider the op a success. Ditto for write_chunk. Fix up eunit test to accomodate the change of semantics.	2015-07-18 23:22:14 +09:00
Scott Lystig Fritchie	87867f8f2e	WIP: bugfix for rare flapping infinite loop (done fix I hope) {sigh} This is a correction to a think-o error in the "WIP: bugfix for rare flapping infinite loop (better fix I hope)" bugfix that I thought I had finished in the slf/chain-manager/cp-mode branch. Silly me, the test for myself as the author of the not_sane transition was wrong: we don't do that kind of insanity, other nodes might, though. ^_^	2015-07-18 17:53:17 +09:00
Scott Lystig Fritchie	19ce841471	Merge slf/chain-manager/cp-mode (fix conflicts)	2015-07-17 16:39:37 +09:00
Scott Lystig Fritchie	b295c7f374	Log more info on private projection write failure	2015-07-17 16:20:54 +09:00
Scott Lystig Fritchie	f4d16881c0	WIP: bugfix for rare flapping infinite loop (better fix I hope) %% So, I'd tried this kind of "if everyone is doing it, then we %% 'agree' and we can do something different" strategy before, %% and it didn't work then. Silly me. Distributed systems %% lesson #823: do not forget the past. In a situation created %% by PULSE, of all=[a,b,c,d,e], b & d & e were scheduled %% completely unfairly. So a & c were the only authors ever to %% suceessfully write a suggested projection to a public store. %% Oops. %% %% So, we're going to keep track in #ch_mgr state for the number %% of times that this insane judgement has happened.	2015-07-17 14:51:39 +09:00
Scott Lystig Fritchie	0a8821a1c6	WIP: bugfix for rare flapping infinite loop (fixed I hope) I'll run a set of PULSE tests (Cmd_e of the 'regression' style) to try to confirm a fix for this pernicious little thing. Final (?) part of the fix: add myself to SeenFlappers in react_to_env_A30().	2015-07-16 23:23:30 +09:00
Scott Lystig Fritchie	b4d9ac5fe0	Hooray, PULSE things look stable; remove debugging verbose cruft	2015-07-16 21:57:34 +09:00
Scott Lystig Fritchie	c10200138c	Hooray??! Fix the damn PULSE hangs by using infinity supervisor shutdown times	2015-07-16 21:17:46 +09:00
Scott Lystig Fritchie	3a4624ab06	Hrm, fewer deadlocks, but lots of !@#$! mystery hangs @ startup & teardown	2015-07-16 20:13:48 +09:00
Scott Lystig Fritchie	d331e09923	Hrm, fewer deadlocks, but sometimes unreliable shutdown	2015-07-16 17:59:02 +09:00
Scott Lystig Fritchie	f2fc5b91c2	Add more PULSE instrumentation -> more deadlocks	2015-07-16 16:25:38 +09:00
Scott Lystig Fritchie	73ac220d75	Add machi_verbose.hrl	2015-07-16 16:01:53 +09:00
Scott Lystig Fritchie	0ead97093b	WIP: bugfix for rare flapping infinite loop (unfinished) part ...	2015-07-16 00:18:42 +09:00
Scott Lystig Fritchie	18c92c98f8	WIP: bugfix for rare flapping infinite loop (unfinished) part IV	2015-07-15 18:42:59 +09:00
Scott Lystig Fritchie	402720d301	WIP: bugfix for rare flapping infinite loop (unfinished) part II	2015-07-15 17:23:17 +09:00
Scott Lystig Fritchie	6f9a603e99	WIP: bugfix for rare flapping infinite loop (unfinished)	2015-07-15 12:44:56 +09:00
Scott Lystig Fritchie	0f667c4356	WIP: add more debugging/react info	2015-07-15 11:25:06 +09:00
Scott Lystig Fritchie	7c970d90a6	Bugfix: use correct updated #state in react_to_env_A30() {sigh}	2015-07-15 00:44:07 +09:00
Scott Lystig Fritchie	5eb6ebc874	Bugfix: add missing remember_partition_hack() calls in perhaps_call path	2015-07-14 17:17:14 +09:00
Scott Lystig Fritchie	fd66fe46b5	Move react logging in react_to_env_A30()	2015-07-14 17:16:23 +09:00
Scott Lystig Fritchie	0089af0a86	Bugfix: moving inner -> outer projection, use calc_projection() for sanity	2015-07-10 21:11:34 +09:00
Scott Lystig Fritchie	f746b75254	Bugfix: A30: if Kicker_p only true if we actually have an inner proj!	2015-07-10 20:25:44 +09:00
Scott Lystig Fritchie	e9e4c54b25	Bugfix: undo the jump directly from A30 -> C100.	2015-07-10 20:24:44 +09:00
Scott Lystig Fritchie	ed7dcd14db	Avoid putting inner_summary in dbg proplist	2015-07-10 17:47:33 +09:00
Scott Lystig Fritchie	4d41c59e19	Bugfix: machi_projection:new/6 derp: argument order mistake	2015-07-10 16:41:28 +09:00
Scott Lystig Fritchie	cf9ae5b555	WIP: correct calc of All_UPI_Repairing_were_unanimous, but now infinite loop in long chains??	2015-07-10 15:30:31 +09:00
Scott Lystig Fritchie	2060b80830	Keep good refactorings from commit a8390ee2 Also, add more misc details to the 'react' breadcrumb trail. Also, save get(react) results into dbg2 whenever we write a private projection, very valuable for debugging. Also: cleanup PULSE code, add regression commands as option and controls with some new environment variables. These regression sequences were responsbile for several fruitful debugging sessions, so we keep them for posterity and for their ability (with new seeds and PULSE) to find new interleavings.	2015-07-10 15:04:50 +09:00
Scott Lystig Fritchie	badcfa3064	Remove comment cruft	2015-07-07 14:32:02 +09:00
Scott Lystig Fritchie	0f3d11e1bf	Bugfix (part II) rare race between just-finished repair and flapping ending The prior commit wasn't sufficient: the range of transitions is wider than assumed by that commit. So, we take one of two options, with a TODO task of researching the other option.	2015-07-07 14:30:21 +09:00
Scott Lystig Fritchie	96ca7b7082	Bugfix for rare race between just-finished repair and flapping ending Fix for today: We are going to game the system. We know that C100 is going to be checking authorship relative to P_current's UPI's tail. Therefore, we're just going to set it here. Why??? Because we have been using this projection safely for the entire flapping period! ... The only other way I see is to allow C100 to carve out an exception if the repair finished PLUS author_server check fails PLUS if we came from here, but that feels a bit fragile to me: if some code factoring happens in projection_transition_is_saneprojection_transition_is_sane() or elsewhere that causes the author_server check to be something-other-than-the-final-thing-checked, then such a refactoring would likely cause an even harder bug to find & fix. Conditions tested: 5 FLUs plus alternating partitions of: [ [{a,b}], [], [{a,b}], [], [{a,b}], [], [{a,b}], [], [{a,b}], [], [{b,a},{d,e}], [{a,b}], [], [{a,b}], [], [{a,b}], [], [{a,b}], [], [{a,b}], [] ].	2015-07-07 01:29:37 +09:00
Scott Lystig Fritchie	54b5014446	WIP: bugfix in transition, just-in-case commit	2015-07-06 23:56:29 +09:00
Scott Lystig Fritchie	9d4b4b1df6	Bugfix: update inner projection based on previous inner projection	2015-07-06 17:38:15 +09:00
Scott Lystig Fritchie	3f8982cbe1	MAJOR WIP: set author's rank to constant 0? Worthwhile??	2015-07-06 16:12:15 +09:00
Scott Lystig Fritchie	471cde1f2c	WIP: debugging fmt shuffle	2015-07-06 16:11:14 +09:00
Scott Lystig Fritchie	8ee3377fa7	Fix a state transition bug (chain manager infinite loop, oops) %% We have a small problem for state transition sanity checking in the %% case where we are flapping and a repair has finished. One of the %% sanity checks in simple_chain_state_transition_is_sane(() is that %% the author of P2 in this case must be the tail of P1's UPI: i.e., %% it's the tail's responsibility to perform repair, therefore the tail %% must damn well be the author of any transition that says a repair %% finished successfully. %% %% The problem is that author_server of the inner projection does not %% reflect the actual author! See the comment with the text %% "The inner projection will have a fake author" in %react_to_env_A30(). %% %% So, there's a special return value that tells us to try to check for %% the correct authorship here.	2015-07-05 14:52:50 +09:00
Scott Lystig Fritchie	920c0fc610	WIP: much better structure for inner projection sanity checking	2015-07-04 16:46:02 +09:00
Scott Lystig Fritchie	8241d1f600	WIP: cruft, needs refactoring	2015-07-04 14:57:38 +09:00
Scott Lystig Fritchie	65ee0c23ec	Adjust author of inner projections to yield same checksum	2015-07-04 01:58:00 +09:00
Scott Lystig Fritchie	cd026303a0	Unused var cleanup	2015-07-04 00:35:05 +09:00
Scott Lystig Fritchie	9b0a5a1dc3	WIP: 1st part of moving old chain state transtion code to new Ha, famous last words, amirite? %% The chain sequence/order checks at the bottom of this function aren't %% as easy-to-read as they ought to be. However, I'm moderately confident %% that it isn't buggy. TODO: refactor them for clarity. So, now machi_chain_manager1:projection_transition_is_sane() is using newer, far less buggy code to make sanity decisions. TODO: Add support for Retrospective mode. TODO is it really needed? Examples of how the old code sucks and the new code sucks less. 138> eqc:quickcheck(eqc:testing_time(10, machi_chain_manager1_test:prop_compare_legacy_with_v2_chain_transition_check(whole))). xxxxxxxxxxxx..x.xxxxxx..x.x....x..xx........................................................Failed! After 69 tests. [a,b,c] {c,[a,b,c],[c,b],b,[b,a],[b,a,c]} Old_res ([335,192,166,160,153,139]): true New_res: false (why line [1936]) Shrinking xxxxxxxxxxxx.xxxxxxx.xxx.xxxxxxxxxxxxxxxxx(3 times) [a,b,c] %% {Author1,UPI1, Repair1,Author2,UPI2, Repair2} %% {c, [a,b,c],[], a, [b,a],[]} Old_res ([338,185,160,153,147]): true New_res: false (why line [1936]) false Old code is wrong: we've swapped order of a & b, which is bad. 139> eqc:quickcheck(eqc:testing_time(10, machi_chain_manager1_test:prop_compare_legacy_with_v2_chain_transition_check(whole))). xxxxxxxxxx..x...xx..........xxx..x..............x......x............................................(x10)...(x1)........Failed! After 120 tests. [b,c,a] {c,[c,a],[c],a,[a,b],[b,a]} Old_res ([335,192,185,160,153,123]): true New_res: false (why line [1936]) Shrinking xx.xxxxxx.x.xxxxxxxx.xxxxxxxxxxx(4 times) [b,a,c] %% {Author1,UPI1,Repair1,Author2,UPI2, Repair2} %% {a, [c], [], c, [c,b],[]} Old_res ([338,185,160,153,147]): true New_res: false (why line [1936]) false Old code is wrong: b wasn't repairing in the previous state. 150> eqc:quickcheck(eqc:testing_time(10, machi_chain_manager1_test:prop_compare_legacy_with_v2_chain_transition_check(whole))). xxxxxxxxxxx....x...xxxxx..xx.....x.......xxx..x.......xxx...................x................x......(x10).....(x1)........xFailed! After 130 tests. [c,a,b] {b,[c],[b,a,c],c,[c,a,b],[b]} Old_res ([335,214,185,160,153,147]): true New_res: false (why line [1936]) Shrinking xxxx.x.xxx.xxxxxxx.xxxxxxxxx(4 times) [c,b,a] %% {Author1,UPI1,Repair1,Author2,UPI2, Repair2} %% {c, [c], [a,b], c, [c,b,a],[]} Old_res ([335,328,185,160,153,111]): true New_res: false (why line [1981,1679]) false Old code is wrong: a & b were repairing but UPI2 has a & b in the wrong order.	2015-07-04 00:32:28 +09:00
Scott Lystig Fritchie	42fb6dd002	WIP: it's clear that the legacy state transition check is broken, II	2015-07-03 23:37:36 +09:00
Scott Lystig Fritchie	caeb322725	WIP: it's clear that the legacy state transition check is broken	2015-07-03 23:17:34 +09:00
Scott Lystig Fritchie	83015c319d	WIP: yeah, now we're going places	2015-07-03 22:05:35 +09:00
Scott Lystig Fritchie	6a706cbfeb	WIP: Refactoring and prototyping goop, broken test	2015-07-03 19:21:41 +09:00
Scott Lystig Fritchie	9b3cd9056a	Un-TEST'ify testr_react_to_env() everywhere	2015-07-03 16:18:40 +09:00
Scott Lystig Fritchie	2b64028bbd	Add kick_projection_reaction, implement yo:tell_author_yo()	2015-07-03 04:30:05 +09:00
Scott Lystig Fritchie	c6870a1c86	If FLU is wedged by a newer client epoch ID, kick the chain manager to react	2015-07-03 02:17:01 +09:00
Scott Lystig Fritchie	ff66638eb3	Sequencer changes file sequence number when epoch_id change is detected	2015-07-03 02:04:04 +09:00
Scott Lystig Fritchie	da3a56dd74	Fix epoch checking in eunit tests and enforcement by FLU (always permit list_files())	2015-07-01 18:12:22 +09:00
Scott Lystig Fritchie	2c869ed598	TODO fix: wedge self	2015-07-01 17:19:11 +09:00
Scott Lystig Fritchie	1e14fe878f	Ha, oops! Add bad_epoch code, derp 1	2015-07-01 15:51:25 +09:00
Scott Lystig Fritchie	a658a64482	Cosmetic formatting change	2015-07-01 15:37:53 +09:00
Scott Lystig Fritchie	a0061d6ffa	make decode_csum_file_entry() very slightly less brittle	2015-07-01 15:18:57 +09:00
Scott Lystig Fritchie	d710d90ea7	Fix usage of checksum_list by machi_chain_repair.erl	2015-07-01 15:04:22 +09:00
Scott Lystig Fritchie	0321e05b46	Fix usage of checksum_list by machi_basho_bench_driver.erl	2015-07-01 15:03:56 +09:00
Scott Lystig Fritchie	e3b80c6ac2	Docuemntation updates	2015-06-30 19:04:23 +09:00
Scott Lystig Fritchie	00c8cf0ef7	Rename temporary HTTP server hack functions	2015-06-30 16:19:44 +09:00
Scott Lystig Fritchie	7542fe8225	WIP: all eunit tests are passing again, yay	2015-06-30 16:12:23 +09:00
Scott Lystig Fritchie	e9d50a2128	WIP: Reinstate one eunit test, fix type bugs	2015-06-30 15:51:03 +09:00
Scott Lystig Fritchie	3d2b49b7e5	WIP: refactoring & edoc'ing	2015-06-30 15:20:35 +09:00
Scott Lystig Fritchie	310fdb1f6a	Add crude file size check to do_server_checksum_listing()	2015-06-30 14:13:26 +09:00
Scott Lystig Fritchie	2d070bf1e3	Minor refactoring + add demo/exploratory time measurement code %% Demo/exploratory hackery to check relative speeds of dealing with %% checksum data in different ways. %% %% Summary: %% %% * Use compact binary encoding, with 1 byte header for entry length. %% * Because the hex-style code is far slower just for enc & dec ops. %% * For 1M entries of enc+dec: 0.215 sec vs. 15.5 sec. %% * File sorter when sorting binaries as-is is only 30-40% slower %% than an in-memory split (of huge binary emulated by file:read_file() %% "big slurp") and sort of the same as-is sortable binaries. %% * File sorter slows by a factor of about 2.5 if {order, fun compare/2} %% function must be used, i.e. because the checksum entry lengths differ. %% * File sorter + {order, fun compare/2} is still far faster than external %% sort by OS X's sort(1) of sortable ASCII hex-style: %% 4.5 sec vs. 21 sec. %% * File sorter {order, fun compare/2} is faster than in-memory sort %% of order-friendly 3-tuple-style: 4.5 sec vs. 15 sec.	2015-06-30 14:08:46 +09:00
Scott Lystig Fritchie	34b046acbd	Remove machi_pb_wrap.erl	2015-06-29 17:31:07 +09:00
Scott Lystig Fritchie	dba7041929	Change names to indicate we're no longer in PB land	2015-06-29 17:20:17 +09:00
Scott Lystig Fritchie	151e696324	WIP: yank out more unused cruft	2015-06-29 17:14:33 +09:00
Scott Lystig Fritchie	87ec988353	WIP: yank out more unused cruft	2015-06-29 17:06:28 +09:00
Scott Lystig Fritchie	6cd3b8d0ec	WIP: yank out lots of unused cruft	2015-06-29 17:02:58 +09:00
Scott Lystig Fritchie	d54c74f58a	WIP: yank out io:format	2015-06-29 16:53:41 +09:00
Scott Lystig Fritchie	7aff9fca70	WIP: giant hairball 12	2015-06-29 16:42:05 +09:00
Scott Lystig Fritchie	b25ab3b7ac	WIP: giant hairball 11	2015-06-29 16:24:57 +09:00
Scott Lystig Fritchie	64817dd7e8	WIP: giant hairball 01	2015-06-29 16:10:43 +09:00

... 2 3 4 5 6 ...

511 commits